Skip to content

Build1 publisherNot yet confirmed elsewhere3 min readPublished

Databricks' Funke parses HL7v2 messages into native Spark columns without converting to FHIR first

Databricks has released Funke, a PySpark library and streaming pipeline that parses HL7v2 messages directly into a native Spark column. Teams that use it to replace a FHIR converter or a flattening vendor take on the gold-layer field mapping themselves, sender quirks included.

The Engineer · Build desk

How we use AISend a correction

Illustration accompanying Databricks' Funke parses HL7v2 messages into native Spark columns without converting to FHIR first
Generated illustration

What happened

  • Funke models each parsed message as a map from segment name to that segment's repetitions, with fields addressed by field number, repetition, component and subcomponent.
  • Databricks says Funke supports every HL7 message type and version, so one pipeline can take the range of versions a health system receives.
  • Funke deploys as a Declarative Automation Bundle, ingests through a Declarative Pipeline and stores everything in Unity Catalog.
  • Funke succeeds Smolder, the Scala Spark data source Databricks open-sourced in 2021 to load HL7v2 messages into DataFrames.

Compiled by The EngineerSomething wrong?How this is made

Why it matters

  • exposure A lossless parse puts every field a sender transmits into the silver table, so access rules on parsed_messages have to cover the full message, beyond the fields any gold table uses.
  • capability Detail that a FHIR mapping would drop stays queryable with ordinary SQL, because the whole hierarchy lands as one Spark column.
  • decision Funke's pipeline starts at files in a Unity Catalog volume, so a team retiring its interface engine still needs something to deliver messages into storage.

Funke's map keys every value by position [4]. That matches how HL7v2 is built, with segments, fields, components and subcomponents separated by characters the message declares in its own header [1]. It also means the parser records where a sender put a value. It has no view of what the sender meant. Databricks' own post says the standard is loose enough that messages in production differ from one sender to another [2]. If two lab systems carry the same unit in different components, Funke keeps both faithfully, at two different addresses. The pipeline defines two tables and leaves the gold tables to the user [9].

The name is German for spark, a nod to its predecessor Smolder [15]; at least the fire is heading the right way. The parsing itself is careful work. It takes the field, component, repetition and subcomponent separators from each message's MSH header and applies the standard escape sequences [5]. Databricks describes the parse as designed to be lossless, with the goal of keeping the original structure as far into the pipeline as possible [7]. The post does not say what the parser does with a message that breaks the encoding rules.

Bronze keeps each decoded message next to an MD5 hash, an insert timestamp and a message ID [6]. Silver adds a single hl7 column holding the parse [10]. Keeping the raw text in bronze is the design choice I like most. If the parser mishandles one sender, I'd expect the fix to be replayable against bronze history without asking anyone to resend a feed.

Databricks frames Funke against two workarounds. Converting to FHIR first adds a translation layer and can drop detail that has no clean FHIR equivalent, the post says [11]. A third-party flattening engine means paying another vendor and moving data off the platform, and its wide tables lose the granular structure underneath [12]. Funke's answer is to postpone the shape decision. The post says users decide which fields matter for each use case instead of accepting the shape a converter or vendor chose [14]. With Funke, that shape is gold-table code the team's own engineers write and maintain [9].

Databricks says a team can go from nothing to a scalable streaming HL7 pipeline in minutes [18]. That figure covers standing up bronze and silver. The per-sender mapping into gold is separate work.

For Funke to replace a flattening vendor in a given shop, two things have to be true. Its parser has to accept every sender's real traffic, including messages that bend the spec. The team also has to own the per-sender mapping that previously lived in the vendor's tables. The first condition can be tested before any migration: replay a month of each feed into silver and count the messages that fail to parse or put a field where the gold code does not expect it.

What to watch

  • Whether Databricks documents how Funke handles malformed messages or segments that fall outside the standard.
  • Reports from health systems that replay multi-sender feeds through parsed_messages and publish their parse-failure rates.
  • Whether Databricks ships reference gold tables for common flows such as admissions and lab results, which would shrink the mapping work teams inherit.

Clarity's read

What the record supports and how the coverage leans. The claims behind it follow.

Reality

Evidence35
Adoption
Insufficient
Hype gap+25
Incentives80
Confidence40
Why these scores

Claim ledger

Ranked by verification strength, evidence, and original report placement.

  1. [1]

    An HL7v2 message is a nested, delimiter-encoded structure: segments made of fields, fields made of components and repetitions, components made of subcomponents, separated by special characters that the message declares in its header.

    ReportedSupportedSource: Databricks blogView cited source
  2. [2]

    The HL7v2 specification leaves room for flexibility, so real-world messages vary from one sending system to the next.

    ReportedSupportedSource: Databricks blogView cited source
  3. [3]

    Funke is a Python and PySpark library plus a ready-to-deploy pipeline that parses an HL7v2 message directly into a native Spark type and keeps the entire hierarchy intact.

    ReportedSupportedSource: Databricks blogView cited source

Sources

1 independent publisher whose own reporting we read for this story.

  1. databricks.com

    1 article · October 8, 2026

    Introducing Funke: Native HL7v2 Parsing on Databricks

Share your take

Let Clarity write the post for you.

Signed-in readers get a short post drafted on this story in the register they choose — narrative, analytical, or a direct position — editable to the last word before it goes anywhere. The share buttons at the top of this story work without an account.

Topics and entities

Follow any of these and your For You feed starts watching them — no settings page required.

Topics

  • Lakehouse data engineeringFollow
  • Healthcare data interoperabilityFollow
Loading related stories