Build1 publisher3 min readPublished
Lookup-field improvement on a low-code platform broke a warehouse app untouched for eight months
One low-code platform's reworked lookup field broke a customer's warehouse app that had run correctly, untouched, for eight months. Writing on dev.to, the team argues each model change is a production data migration on apps it cannot test.
The Engineer · Build desk
What happened
- By the Thursday afternoon after the release, the customer's warehouse app had stopped assigning storage locations.
- The person who built the app had left the customer's company in March.
- The team's release notes were accurate and its tests passed, while the customer's app had been built against the old lookup behavior.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- constraint The platform cannot list the customer workflows a field change will break before it ships, because those references live in user-written rows that no compiler reads.
- cost Letting customers choose their upgrade date commits the platform to maintaining old behavior for as long as the slowest customer waits, and those tend to be the most fragile apps.
- exposure Apps whose authors have moved on still take every release, so a behavior change can reach live operations with nobody on the customer side who knows how the app works.
The post states the constraint in five words: "the application model is data." [6] A customer's app on the platform is tables, fields, validation rules, workflow steps and permission trees, stored as rows that users wrote and other rows reference [9]. Every release is deployed to every one of those apps, including, in the post's example, one a contractor assembled in 2022 and abandoned [9]. So every change to the model is a data migration on someone else's production data, typically with no rollback path [7].
That changes what a rename costs. In a codebase, the compiler finds every reference that broke. On the platform, the references sit in customer rows. "Rename a field on someone else's application: nothing tells you anything, and the workflow that referenced it fails silently at 2 a.m. on a Saturday," the author wrote [8].
The warehouse case was harder to catch than a rename. Of the three changes in the release, the permission check moved earlier in the write path looks like the risky one on paper [1]. The break came from the lookup field, whose new behavior was better in every case the team had tested [4]. The customer's app had been written against the old behavior, "deliberately or not," and had run correctly for eight months [5]. A green suite is a claim about the vendor's own cases. For it to cover this one, the suite would have to contain the customer's usage, and on this platform that usage exists only as the customer's metadata [9]. The post puts it as a blast radius measured in somebody's business process, with no test suite for it [14].
The post compares rollout strategies by who absorbs the cost. Shipping everything at once is fast for the platform and brutal for customers, and a bad fix cannot be un-shipped [10]. "The velocity you gain is borrowed against trust you will need later," the author wrote [10]. A flag day lets each customer choose when to take new behavior. The platform then supports the old path for as long as the slowest customer delays [11]. The post adds that "the customers who delay longest are, almost by definition, the ones whose applications are the most fragile and least understood." [11]
I think the author lands on the right tradeoff for a platform whose customers run operations on it. "That was the week I stopped thinking about upgrades as a release-engineering problem, and started thinking about them as the hardest product decision a low-code platform makes," the author wrote [13]. The post now calls backward compatibility "the primary surface of the product" [12]. In my view, the incident supports treating a changed default on an existing field type as a breaking change, even when the new behavior tests better.
The evidence is one incident, told by the vendor two weeks after the release [15]. The post does not name the platform or say how many applications took the release. It establishes that a tested, accurately documented release can break a stable app, with no measure of how often that happens.
What to watch
- The third rollout strategy the post counts, and whether it tests new releases against customers' stored app metadata before rollout.
- Any count from the vendor of how many applications the lookup-field change touched, and whether others broke besides the warehouse app.
- Whether the platform starts versioning field-type behavior per application, so an old app keeps the lookup semantics it was built against.