Build1 publisherNot yet confirmed elsewhere3 min readPublished
Polars 2.0's streaming default can reorder group_by and join output without an error
Polars 2.0 runs LazyFrame.collect() on its streaming engine by default, a dev.to upgrade guide reports. One environment variable pins the old engine, so teams can fix the loud API breaks before hunting row-order changes that raise no error.
The Engineer · Build desk
What happened
- The dev.to guide counts eight changes in the official Polars 2.0 upgrade notes that can alter a pipeline's output without raising an error.
- Under streaming, unpivot, group_by and joins no longer guarantee row order, so steps that take a group's first row or zip frames by position can change between runs.
- Horizontal pl.concat now raises when frame heights differ, and the old null-padding behaviour moves to how="horizontal_extend".
- read_csv now runs as scan_csv().collect(), drops n_threads, batch_size, sample_size and rechunk, and takes schema_overrides in place of dtypes.
- As of 8 October 2026, sources disagreed on whether Polars 2.0.0 final or only a release candidate had shipped.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Setting the environment-variable pin turns a 2.0 upgrade into two separate test runs, so each failure points at either the API changes or the engine switch.
- exposure Reports built on a group's first row, positional zips or hand-diffed CSV exports can ship different numbers between runs while suites without row-order checks stay green.
- constraint Pinning the old engine buys time on row order alone; tightened casts, the concat height check and the new selector meaning arrive with the version bump itself.
Code that never passed an engine argument is running engine="auto". In 2.0 that setting resolves to streaming for both LazyFrame.collect() and collect_async() [1]. Eager DataFrame code does not move [3]. According to the dev.to guide, streaming uses less memory on large data and executes differently [15].
The guide's advice is to pin the in-memory engine, run the tests, then remove the pin one pipeline at a time [16]. There are three places to pin: engine="in-memory" on a single call, pl.Config.set_engine_affinity("in-memory") in code, or POLARS_ENGINE_AFFINITY=in-memory in the environment [4]. I'd start with the environment variable, as the guide does. It changes no source file, so the whole suite runs against 2.0 before anyone edits code [5]. In a repository where CI runs every pipeline's tests, that splits the upgrade into two experiments. Failures with the pin set come from API changes. The guide says those fail fast and are easy to fix [19]. Anything that breaks after a pipeline is unpinned comes from the engine [20]. Deployment config needs one edit either way: POLARS_STREAMING_CHUNK_SIZE is now POLARS_IDEAL_MORSEL_SIZE [6].
The pin covers the engine behind collect() and nothing else [21]. Casts tighten regardless. Integer to Categorical and back is disallowed, as is String to Date, Datetime or Time, and the replacement is the explicit parser [13]. A list passed to schema_overrides must now cover every column [11]. The guide also says a selector expression that used to mean a set intersection now means something else [14]. Running on the old engine does not bring the old meaning back [21].
For row order the guide offers three fixes: maintain_order=True on group_by, maintain_order="left" on a left join, or a sort at the end of the query [8]. It prefers the final sort as cheaper to reason about than preserving order inside every operation [17]. For report pipelines I agree. One sort at the output boundary is one line to review, while maintain_order is a flag someone has to remember on each operation. The guide's list of exposed steps includes a CSV that a human diffed [7]. That reviewer is an assertion library with no version pin. The guide also says to keep check_row_order on in assert_frame_equal and let the failures show where order mattered [9]. Streaming output can differ between runs [7], so a single green run with that check on proves little, and I would repeat the comparison before removing a pin.
Horizontal concat is the change both safeguards can miss. The engine pin does not touch it [21]. The guide says it is loud only when tests use frames of uneven height [18]. A suite whose fixtures always had matching heights passes, and the pipeline fails in production the day an upstream filter returns fewer rows [18].
What to watch
- The 2.0.0 final release notes on GitHub, and whether they match the guide's list of silent changes, including the exact new meaning of the changed selector.
- Whether a later 2.x release guarantees row order for group_by and joins under the streaming engine.
Clarity's read
What the record supports and how the coverage leans. The claims behind it follow.
Reality
- Evidence50
- Adoption
- Insufficient
- Hype gap+5
- Incentives20
- Confidence50
Claim ledger
Ranked by verification strength, evidence, and original report placement.
- [1]
In Polars 2.0, with engine="auto", LazyFrame.collect() and collect_async() pick the streaming engine; the Polars 1.x default was in-memory.
- [2]
Eight changes in the official Polars 2.0 upgrade guide can alter pipeline output without raising a single error.
- [3]
Eager DataFrame operations are unchanged in Polars 2.0, so code written in the eager style does not move.
- [4]
Three switches restore the in-memory engine: lf.collect(engine="in-memory") per call, pl.Config.set_engine_affinity("in-memory") globally in code, or the environment variable POLARS_ENGINE_AFFINITY=in-memory.
- [5]
The environment variable is the safest first step because it changes no source file, so the whole test suite can run against 2.0 before touching code.
- [6]
The environment variable POLARS_STREAMING_CHUNK_SIZE is now POLARS_IDEAL_MORSEL_SIZE, so deployment config that sets it needs updating.
- [7]
The streaming engine does not guarantee row order for unpivot, group_by and joins; if a downstream step took the first row of a group, wrote a CSV that a human diffed, or zipped two frames by position, output can differ run to run.
- [8]
Fixes for row order: group_by(..., maintain_order=True), a sort at the end of the query, or join(..., how="left", maintain_order="left") to preserve the left frame's order.
- [9]
Tests that compare frames with assert_frame_equal should keep check_row_order on and let the failures show where order mattered.
- [10]
read_csv is now scan_csv().collect(); it loses n_threads, batch_size, sample_size and rechunk, and the dtypes argument is now schema_overrides.
- [11]
A list passed to schema_overrides must cover every column.
- [12]
pl.concat(frames, how="horizontal") used to pad shorter frames with nulls; in 2.0 it raises when heights differ, and the padding behaviour exists under how="horizontal_extend".
- [13]
Casting integers to Categorical and Categorical to integers is disallowed, casting String to Date, Datetime or Time is disallowed, and the replacement is the explicit parser.
- [14]
A selector expression that used to mean a set intersection now means something else in Polars 2.0.
- [15]
The streaming engine uses less memory on large data and produces different execution behaviour.
- [16]
Recommended path: pin the in-memory engine first, run the tests, then remove the pin one pipeline at a time.
- [17]
Sorting at the end is cheaper to reason about than preserving order inside every operation.
- [18]
The horizontal concat change is a loud break only if tests use uneven frames; in a pipeline that always had matching heights it passes silently, then fails in production the day an upstream filter returns fewer rows.
- [19]
The loud breaks, such as removed methods and renamed arguments, fail fast and are easy to fix.
- [20]
With the environment-variable pin set, test failures on 2.0 come from API changes; failures that appear after a pipeline is unpinned come from the engine switch.
- [21]
The engine pin only selects the engine behind collect(); the cast, schema_overrides, selector and horizontal concat changes apply in 2.0 whether or not it is set.
- [22]
As of 8 October 2026 sources disagree on whether Polars 2.0.0 final has shipped or only a release candidate; the guide advises checking the GitHub releases page before pinning a version.
Sources
1 independent publisher whose own reporting we read for this story.
- dev.toPolars 2.0 Upgrade Guide: 8 Breaking Changes That Hit Silently
1 article · October 8, 2026
Topics and entities
Follow any of these and your For You feed starts watching them — no settings page required.