Build1 publisher3 min readPublished
A .keras config can carry a marshalled Python code object, and load_model runs it
"Use .h5 or .keras instead of pickle" is not a mitigation. Scanning pipelines built on the pickle-only model of executable artifacts have a hole in them.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction
What happened
- Keras lets you define a layer as an arbitrary Python callable via Lambda.
- Because the callable must survive serialisation, it is stored in the Keras model config as a marshalled code object.
- When load_model reconstructs the model, it runs the stored code object.
- This is not an exploit; it is the documented behaviour of a feature people use.
- A .keras file is a code-carrying format in exactly the way a pickle is, and a pipeline that treats .h5 as the safe alternative to pickle is mistaken.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
A post from the AIsbom project, published on dev.to, sets out a detail that undercuts the standard remediation for malicious model files: a Keras model config can contain a marshalled Python code object, and `load_model` executes it while reconstructing the model [1][2][3]. That matters because the usual advice when a team finds pickles in `models/` is to switch to `.h5` or `.keras`, and according to the post a pipeline that treats `.h5` as the safe alternative to pickle is mistaken [5].
The mechanism is not an exploit. Keras lets you define a layer as an arbitrary Python callable via `Lambda`; the callable has to survive serialisation, so it is stored in the config as a marshalled code object [1][2]. The post is explicit that this is documented behaviour of a feature people use [4]. The consequence for a scanner is unchanged by that: a `.keras` file hands the loader something executable in exactly the way a pickle does [5].
There is a second-order trap for anyone writing the detector. The obvious way to inspect a marshalled blob is to unmarshal it, and `marshal.loads` on untrusted input is itself unsafe [6]. The post's answer is that the blob's type can be read from its header bytes instead, and that in AIsbom v1.3.0 Lambda layers and embedded code objects are flagged CRITICAL with the payload never unmarshalled [7][8]. Both containers Keras writes, the `.keras` zip and legacy HDF5, are handled without pulling an HDF5 library into the install [9]. Those are the project's claims about its own release, not independent test results.
The same shape recurs in two other formats the pickle-only mental model waves through. ONNX has a reputation as safe because there is no embedded bytecode, which the post calls mostly true with two caveats: the external-data mechanism stores a path that the loader reads, so a path pointing outside the model directory turns loading into an arbitrary-file read, and operators can come from a non-standard domain, meaning the graph depends on something other than the standard operator set to execute [10][11][12]. Neither needs the graph run to detect; you walk the protobuf, including the subgraphs carried by `If`, `Loop` and `Scan` [13]. GGUF files can embed a `chat_template`, a Jinja template that downstream code renders to format prompts, and Jinja's sandbox has known escape constructs, so the check has to be static because rendering is the vulnerability [14]. That is three non-pickle formats where the loader or the analysis step does the attacker's work [22].
Even inside pickle, the post lists implementations that look plausible and are wrong: stopping at the first `STOP` opcode, when a legacy `torch.save` file hides its object behind several header pickles [15]; trusting the container, when the same checkpoint content packed as 7z leaves a ZIP-only scanner reporting nothing [16]; skipping corrupt files, when the pickle VM runs sequentially and a payload at the front executes before the damaged tail is reached [17]; and sniffing file type before disassembly, when a protocol-0 pickle is printable ASCII and can pass as a text config [18].
What to watch: whether your scanner reports an unfinished scan as unfinished. AIsbom added two MEDIUM levels for that, one for hitting a work limit and one for an archive member it cannot read at all, on the grounds that a loader without integrity checks would still run that member [20]. The post's line is that a scanner which quietly gives up and returns clean is worse than none, because you act on the clean bill [21]. Also worth checking in your own allowlists: whether a global is judged by its resolved module and attribute, so a submodule does not inherit an allowlisted parent's trust [19].