Build1 publisherNot yet confirmed elsewhere3 min readPublished
Databricks' Funke parses HL7v2 messages into native Spark columns without converting to FHIR first
Databricks has released Funke, a PySpark library and streaming pipeline that parses HL7v2 messages directly into a native Spark column. Teams that use it to replace a FHIR converter or a flattening vendor take on the gold-layer field mapping themselves, sender quirks included.
The Engineer · Build desk

What happened
- Funke models each parsed message as a map from segment name to that segment's repetitions, with fields addressed by field number, repetition, component and subcomponent.
- Databricks says Funke supports every HL7 message type and version, so one pipeline can take the range of versions a health system receives.
- Funke deploys as a Declarative Automation Bundle, ingests through a Declarative Pipeline and stores everything in Unity Catalog.
- Funke succeeds Smolder, the Scala Spark data source Databricks open-sourced in 2021 to load HL7v2 messages into DataFrames.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- exposure A lossless parse puts every field a sender transmits into the silver table, so access rules on parsed_messages have to cover the full message, beyond the fields any gold table uses.
- capability Detail that a FHIR mapping would drop stays queryable with ordinary SQL, because the whole hierarchy lands as one Spark column.
- decision Funke's pipeline starts at files in a Unity Catalog volume, so a team retiring its interface engine still needs something to deliver messages into storage.
Funke's map keys every value by position [4]. That matches how HL7v2 is built, with segments, fields, components and subcomponents separated by characters the message declares in its own header [1]. It also means the parser records where a sender put a value. It has no view of what the sender meant. Databricks' own post says the standard is loose enough that messages in production differ from one sender to another [2]. If two lab systems carry the same unit in different components, Funke keeps both faithfully, at two different addresses. The pipeline defines two tables and leaves the gold tables to the user [9].
The name is German for spark, a nod to its predecessor Smolder [15]; at least the fire is heading the right way. The parsing itself is careful work. It takes the field, component, repetition and subcomponent separators from each message's MSH header and applies the standard escape sequences [5]. Databricks describes the parse as designed to be lossless, with the goal of keeping the original structure as far into the pipeline as possible [7]. The post does not say what the parser does with a message that breaks the encoding rules.
Bronze keeps each decoded message next to an MD5 hash, an insert timestamp and a message ID [6]. Silver adds a single hl7 column holding the parse [10]. Keeping the raw text in bronze is the design choice I like most. If the parser mishandles one sender, I'd expect the fix to be replayable against bronze history without asking anyone to resend a feed.
Databricks frames Funke against two workarounds. Converting to FHIR first adds a translation layer and can drop detail that has no clean FHIR equivalent, the post says [11]. A third-party flattening engine means paying another vendor and moving data off the platform, and its wide tables lose the granular structure underneath [12]. Funke's answer is to postpone the shape decision. The post says users decide which fields matter for each use case instead of accepting the shape a converter or vendor chose [14]. With Funke, that shape is gold-table code the team's own engineers write and maintain [9].
Databricks says a team can go from nothing to a scalable streaming HL7 pipeline in minutes [18]. That figure covers standing up bronze and silver. The per-sender mapping into gold is separate work.
For Funke to replace a flattening vendor in a given shop, two things have to be true. Its parser has to accept every sender's real traffic, including messages that bend the spec. The team also has to own the per-sender mapping that previously lived in the vendor's tables. The first condition can be tested before any migration: replay a month of each feed into silver and count the messages that fail to parse or put a field where the gold code does not expect it.
What to watch
- Whether Databricks documents how Funke handles malformed messages or segments that fall outside the standard.
- Reports from health systems that replay multi-sender feeds through parsed_messages and publish their parse-failure rates.
- Whether Databricks ships reference gold tables for common flows such as admissions and lab results, which would shrink the mapping work teams inherit.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence35
- Adoption
- Insufficient
- Hype gap+25
- Incentives80
- Confidence40
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
An HL7v2 message is a nested, delimiter-encoded structure: segments made of fields, fields made of components and repetitions, components made of subcomponents, separated by special characters that the message declares in its header.
- [2]
The HL7v2 specification leaves room for flexibility, so real-world messages vary from one sending system to the next.
- [3]
Funke is a Python and PySpark library plus a ready-to-deploy pipeline that parses an HL7v2 message directly into a native Spark type and keeps the entire hierarchy intact.
- [4]
A parsed message is modeled as a map from segment name to the repetitions of that segment; each field is itself a map, addressable by field number, then repetition, then component, then subcomponent.
- [5]
The parser handles the field, component, repetition and subcomponent separators declared in the MSH header, plus the standard escape sequences.
- [6]
New HL7 files arrive in a Unity Catalog volume; Auto Loader picks them up, decodes the content and writes them to the raw_messages table alongside ingestion metadata including an MD5 hash for tracking a message through the pipeline, an insert timestamp and a message ID.
- [7]
The parsing is designed to be lossless, with the goal to retain all of the original structure as far into the pipeline as possible.
- [8]
In 2021 Databricks open-sourced Smolder, a Scala Spark data source that loaded HL7v2 messages into DataFrames; Funke is its successor.
- [9]
Funke deploys as a Declarative Pipeline following the medallion pattern; it defines two tables, and users build gold tables on top for their own use cases.
- [10]
The parsed_messages table reads the raw stream and applies Funke's parser, adding a single hl7 column of the native type.
- [11]
According to Databricks, converting HL7v2 to FHIR first adds a translation layer and can drop detail that never had a clean FHIR equivalent.
- [12]
According to Databricks, handing messages to a third-party engine that flattens the hierarchy into wide tables means paying another vendor, moving data out of the platform, and losing direct access to the granular structure underneath.
- [13]
Funke deploys as a Declarative Automation Bundle (DAB), ingests through a Declarative Pipeline, and stores everything in Unity Catalog.
- [14]
Databricks says users get to decide which fields matter for a given use case, rather than accepting whatever shape a converter or a vendor chose for them.
- [15]
The name Funke is German for spark, a nod to its lineage from Smolder.
- [16]
Because the parsed message is a normal Spark column, every element is reachable with ordinary DataFrame or SQL expressions.
- [17]
Databricks says Funke supports every HL7 message type and version, so the same pipeline can accept messages across the range of versions a health system typically receives.
- [18]
Databricks says users can go from nothing to a scalable, streaming HL7 pipeline in minutes.
Sources
1 independent publisher whose own reporting we read for this story.
- databricks.comIntroducing Funke: Native HL7v2 Parsing on Databricks
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.
Entities
- DatabricksFollow
- FunkeFollow
- SmolderFollow
- HL7 Version 2Follow
- FHIRFollow
- Apache SparkFollow
- Unity CatalogFollow