Build1 publisher3 min readPublished Updated
Inference inside the SELECT: the point is the governance boundary, not the syntax
Databricks is pitching SQL-callable AI functions as convenience. The consequential change is that predictions stop leaving the catalog, and the audit trail stops being a promise.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Databricks describes the pre-existing workflow: an analyst who wants sentiment on support tickets has to ship the rows out to a service, wait for predictions, and stitch them back into a table by hand; it is slow, it breaks when a schema changes, and it introduces unnecessary security and governance risks.
- Databricks says AI Functions bring the AI to the data rather than moving data to a separate AI environment: models are invoked within standard SQL queries, keeping the entire inference process within existing pipelines and Unity Catalog governance.
- Databricks says it manages planning, parallelization and retries so users do not have to handle cluster management or external orchestration.
- Databricks says it is just as easy to run inference on millions of rows as on one row, and the same query scales without rewriting.
- AI function usage appears in system.billing.usage alongside standard Databricks SQL warehouse costs.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
Databricks has published a walkthrough of calling model inference from inside SQL in its warehouse, using task-specific functions including `ai_classify`, `ai_extract`, `ai_translate` and `ai_parse_document` [8]. The interesting part is not that a `SELECT` can now return a sentiment label; it is where the governance boundary sits when it does [4].
The company's own description of the status quo is the useful bit. An analyst who wants sentiment on support tickets ships the rows out to a service, waits for predictions, and stitches them back into a table by hand, which is slow, breaks when a schema changes, and introduces security and governance risk [3]. Every one of those hops is an export. Rows leave the governed table, land in some service's request logs, and return as a join key that a person now maintains. The control story for that path is a set of assurances, not an enforcement point.
Running the model call inside the query plan collapses that. Databricks says inference stays within existing pipelines and Unity Catalog governance [4], and that in the document case lineage runs from the raw PDF to the extracted rows inside a single query plan [12]. In their invoice example, `ai_parse_document` is pointed at a Databricks volume, emits JSON, and passes it to `ai_extract`, which names the entities to pull out, producing a structured table [10][11]. The Python OCR service, the LLM call and the JSON-flattening step that teams normally build by hand fold into the query [13]. Databricks claims this removes the need for fragile custom OCR pipelines or third-party parsers that break on schema changes [18], which is a vendor claim about vendor code and should be read as one.
The line that operators should care about most is the least glamorous: AI function usage lands in `system.billing.usage` alongside standard Databricks SQL warehouse costs [7]. That is the same shift as the governance argument, applied to money. Inference spend becomes queryable by the team that already queries warehouse spend, rather than arriving as a separate invoice someone reconciles at quarter end.
Two limits are worth stating plainly. First, the boundary only extends as far as the catalog does; the argument holds if Unity Catalog is the system of record for the data in question [4], and not otherwise. Second, the post carries no accuracy, latency or per-row cost figures, and no treatment of model version pinning or output drift [19]. `ai_classify` is zero-shot mapping of free text into user-defined labels with no training step [14], which moves label quality into prompt and label design, where regressions are quiet. Databricks also says it handles planning, parallelization and retries, and that a query over millions of rows scales the same as one row without a rewrite [5][6]; that is exactly the kind of claim to test on your own skew before believing.
Worth watching: whether the billing records are granular enough to attribute cost per query and per function rather than per warehouse [7]; whether the narrow functions stay materially cheaper than general-purpose inference, which is the stated rationale for having four of them [8][20]; and whether teams call these from notebooks, Lakeflow Spark Declarative Pipelines and Workflows [9], which puts the calls under job orchestration instead of leaving them in ad hoc analyst SQL.