Build1 distinct publisher3 min readPublished
Append, AUTO CDC and REPLACE WHERE flows can now be declared in plain SQL. Scheduling and state move into the platform; the sequencing decision moves to the analyst.
The Engineer · Build desk

Compiled by The EngineerSomething wrong?How this is made
The unit being shipped here is a standing object that holds state, not a statement that runs once and finishes. Databricks says an APPEND flow tracks new versus previously processed data in the source and manages the underlying serverless pipeline on its own [8]. That is the trade in one sentence: state that used to live in a checkpoint path or somebody's scheduler moves inside the platform, where it is not in a repo and not in a review.
Change data capture is where that matters. Databricks' own framing is that teams reach for MERGE INTO, then discover CDC records arriving out of order and needing extra logic to avoid incorrect results [9]. AUTO CDC replaces the hand-written version with declared keys, a sequencing specification, delete handling, and a choice of SCD Type 1 or Type 2 [10]. The difficulty did not disappear, it changed hands. Picking the sequencing column is the correctness decision in a CDC pipeline, and it is now made by whoever types the flow, in a surface built for querying, with no merge code left for a reviewer to argue with.
The named reference is making an argument about decomposition rather than about SQL. Adrien Marteau of bsport describes decoupling table loads from a single pipeline, processing third-party data independently, and getting better failure management and freshness out of it [11]. That benefit comes from breaking one pipeline into per-source units. It does not require those units to be written in the SQL editor; it does require them to be cheap enough to create that you make a lot of them.
Read the three patterns as a set and they are not equivalent. Append and REPLACE WHERE take away work that was mostly wasted: repeated insert logic with manual scheduling around it [4], and expensive full recomputes standing in for a partition or date-range overwrite [12]. Handing those to the platform costs a team very little judgment. CDC is the one where the hand-coded version carried judgment, and it is now the one-liner.
Two things the material does not settle. Only the append pattern is described as bringing its own automatically managed serverless pipeline; the CDC and REPLACE WHERE descriptions do not say the same [3]. And there is no pricing, compute, or availability detail for any of the three [2]. A flow declared in a query editor that then refreshes on a schedule indefinitely [7] is a recurring bill, and the announcement names no owner for it.
Worth keeping the scale honest: Databricks says thousands of SQL-first users already lean on Materialized Views and Streaming Tables [3], and that these patterns already existed as Lakeflow APIs [6]. The primitives are not the news. The authoring surface, and therefore the job title of the person who now operates them, is.
Ranked by verification strength, evidence, and original report placement.
Databricks is bringing declarative ETL to data warehousing workflows in Lakehouse so SQL users can define common ETL patterns directly within their SQL queries, instead of working in a dedicated pipelines-oriented environment.
Databricks describes this as part of a broader strategy to bring the declarative execution model behind Apache Spark Declarative Pipelines into more authoring experiences across Databricks.
Databricks says thousands of SQL-first users already rely on declarative primitives such as Materialized Views and Streaming Tables.
Databricks says many recurring ETL patterns are easy to describe but difficult to operate, requiring custom SQL logic, manual scheduling and orchestration glue, including appending new records, applying CDC changes, and refreshing only changed data.
The first declarative operations available in the Lakehouse SQL editor map to three recurring ETL patterns: append-only updates, change data capture, and batch overwrites.
Databricks says many of these patterns are already available via declarative APIs such as AUTO CDC in Lakeflow, and that the change is making them accessible directly in Lakehouse for SQL analysts.
Follow any of these and your For You feed starts watching them — no settings page required.
Evidence-backed comparisons of source perspectives and observed adoption signals. Read the methodology
Which Builder, Operator, and Investor concerns the observed source mix emphasized—not a truth score.
Evidence, demonstrated adoption, hype gap, incentives, and confidence are assessed independently, each on its own current evidence. How these are measured.
Detailed but entirely first-party
The cluster is one vendor product blog. It is specific and internally consistent about mechanics - the three flow types, the AUTO CDC parameter surface, predicate-scoped REPLACE WHERE, trigger and orchestration paths - which is more than marketing abstraction. But nothing is independently corroborated: no second publisher, no third-party test, no documentation of release stage, and the one quantified performance figure is vendor-run. Two derived findings mark the gaps: no pricing or availability status, and managed-serverless-execution language that appears only for the append flow.
One named customer, baseline usage unquantified
Adoption evidence for the new SQL editor flows is thin: a single named production user (bsport) speaking qualitatively about SQL AUTO CDC, plus a 'thousands of SQL-first users' figure that describes the older Materialized Views and Streaming Tables primitives rather than the newly announced flows. No release stage, region coverage, or customer counts are given for APPEND, AUTO CDC, or REPLACE WHERE in the SQL editor, so measured adoption stays low even though the underlying declarative engine clearly has a user base.
Moderately overstated by vendor framing
The capability claims themselves are plausible and narrowly scoped - this is largely re-surfacing existing Lakeflow declarative APIs to SQL analysts, which the blog concedes. The overstatement is in the supporting evidence and the elisions: an unaudited 3.4x/2.5x benchmark presented without workload detail, a single testimonial standing in for adoption, silence on release stage and cost, and no acknowledgement that CDC correctness responsibility shifts to whoever types the sequencing parameter. Positive but modest, not a case of an unbuilt product being sold.
Vendor announcing and benchmarking its own product
Every element of the cluster originates with Databricks: the product framing, the usage figure, the customer quote it solicited, and the benchmark it ran comparing its own incrementalization engine against a baseline it defines. The commercial interest is direct - widening the declarative surface pulls warehouse and BI workloads onto Databricks-managed state, scheduling, and serverless compute. There is no adversarial or independent voice in the cluster to offset this.
Confident on what shipped, not on impact
Confidence is moderate. What the product does is stated first-hand by the party that built it, at enough parameter-level detail to be actionable, so the capability claims are reasonably reliable. Everything downstream - performance, cost, adoption breadth, readiness for production rollout - rests on unverified vendor assertion from a single publisher, and two explicit disclosure gaps (availability/pricing, execution substrate for two of three flows) remain open.
invest
Databricks raises $5B at $190B, and the multiple barely moved2 distinct publishers
product
A DOJ probe of a16z board seats asks whether venture portfolios are interlocking directorates1 distinct publisher
invest
Airwallex marks itself up 37% in six months, and tells you why it is not listing1 distinct publisher
build
Databricks quietly switched on dormant MANAGE grants. Check who just became an admin.1 distinct publisher
Distinct publishers with included, body-backed reporting in this cluster.
1 article · August 25, 2026