I'm not following the trends closely, but has Polars become a full replacement for Pandas? Are there use cases where one is better suited than the other?
Polars is effectively a full replacement for Pandas for 99.9% of all cases. The only exception I'm really aware of is if you're working with geospatial data, as there isn't yet a "Geopolars" equivalent of the commonly used "Geopandas". However, Geopolars is still in active development and should eventually be production ready.
Imitation is the sincerest form of flattery! GeoPandas — and its underlying libraries of shapely and GEOS — is an incredible production-ready tool.
GeoPolars is nowhere near the functionality or stability of GeoPandas, but competition is good and, due to its pure-Rust core, GeoPolars will be much easier to use in WebAssembly.
Yes and no, its not replacing the reason why pandas was popular ie data scientists, but it a full replacement of its pipeline usage, And I would saw also beating out spark
my understanding is Polars is faster, scales better without using external solutions, better API, +Rust. Pandas wins if you want to use what the vast majority of folks are using and have used in the past. Probably has a more complete set of helpers / recipes for the little things you bump into when using it thoroughly, but in the age of LLMs, I think that's minor.
>Pandas wins if you want to use what the vast majority of folks are using
Vast majority of skilled developers are now using Polars, unless they are constrained by lack of Narwhals support in their third-party library of choice (e.g. Great Expectations, SHAP). That's the more important trend to follow.
Learn SQL and interface with Duck. You will be 100x faster than Pandas/Polaris duo at fraction of memory. Also SQL is supported literally everywhere with a much more capable than Pandas API. Duck outputs to a Pandas Dataframe, but just treat that like a dictionary. Do all your processing, filtering and aggregation in Duck.
Also, try fx.wtf as a replacement for jq. it comes with a in-built tui viewer that supports vi-keybindings. Ecmascript is built into fx.wtf so you can query the JSON with JS notation (where JSON was born). You can use any JS functions including map/reduce/filter or perform any kind of transformation instead of learning jq dsl that you will forget tomorrow.
There are awkward things. For example, if you ingest a nanosecond resolution timestamp, there's no way to re-export that out of the Polars dataframe with nanosecond resolution.
A bit surprised about the datafusion results from the post, I have tried it time and time again, but datafusion has always been the leading/trading blowers with polars for our workfloads with duckdb being vastly slower.
Doing the benchmarks for 2.0 on the large AWS metal machines at small data sizes (SF=10) really opened my eyes that we have some low-hanging fruit in Polars when it comes to optimizing our constant overhead for smaller queries.
For example our join currently does a full partition into T partitions, for each of the T threads. Overall we create T^2 partitions, which on a 192-core machine is non-trivial. Great if you have a ton of data to feed that with, but if you 'only' have a few dozen million rows it becomes rather small. This is the primary reason we saw in the benchmarks that Polars pinned to 32 threads beats 192 thread Polars at SF=10.
I'll be working on improving that soon. I expect that to have a big impact on SF=10, and a decent impact on ClickBench, which sits between SF=10 and SF=100 in terms of rows.
Although I find pandas a bit aggravating in many ways, for myself and my equally idiotic laboratory scientist pals, seems that it is the default way you might interface with other libraries like SciPy (i.e. they expect things as NumPy arrays or pandas dataframes). Is this a real issue or will most things happily accept a polars dataframe? We don’t work with such large datasets that speed is likely a huge concern tbh.
Actually yes. We had already been planning to do 2.0 for a long time. We originally said we'd move on from 1.x quickly when released 1.0 but ended up staying at 1.x much longer than intended.
From a quick check our first PRs were merged to the 2.0 branch in June:
2026-06-17T21:27:51Z #27993 chore: Stop coercing `pl.col(...)` to selector ...
2026-06-18T14:19:26Z #27996 chore!: Replace multi-seed hash API with a single seed
2026-06-19T07:05:11Z #27991 chore(python!): Remove `Expr.flatten` function
Thing I care about most is whether the old eager-vs-lazy footguns got cleaned up. Half my bugs were a stray collect() in a loop killing the query plan.
I can use pandas to clean a dataset, but each cleaning task is usually one line of code. OTOH, With DuckDB with one SQL statement I can replace 40+ lines of polars/pandas. You may reply, SQL isn't as easy to understand! Fair point, it's a declarative language... which is why I use Malloy. Malloy is to TypeScript as Javascript is to SQL. Malloy is much easier to read and write (just as TypeScript is) because it has a built in semantic model -- all the joins, measures, and dimensions are done in one place.
Here is an example [1] of visualizing college football games. Here are all the queries, and semantic model that power all the visualizations [2] Here is the AI generated typescript/react that does the visualizations [3]. The Malloy ecosystem has Malloyyo and Publisher which are replacements for PowerBI and Tableau and Looker. Here is another example for visualizing global trade [4].
I'm not following the trends closely, but has Polars become a full replacement for Pandas? Are there use cases where one is better suited than the other?
From recent Python Bytes podcast (https://pythonbytes.fm/episodes/show/496/a-lake-house-in-sea...)
> 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory
Python Vs Rust : In terms for speed - No comparison
(The above episode transcript has a link to blog post titled "Pandas should go extinct" )
Polars is effectively a full replacement for Pandas for 99.9% of all cases. The only exception I'm really aware of is if you're working with geospatial data, as there isn't yet a "Geopolars" equivalent of the commonly used "Geopandas". However, Geopolars is still in active development and should eventually be production ready.
it’s on the correct path. i use rust for geo spatial and the gap with c, c++ closing rapidly or negligible in most cases
from https://github.com/pola-rs/geopolars/tree/main
Comparison with GeoPandas
Imitation is the sincerest form of flattery! GeoPandas — and its underlying libraries of shapely and GEOS — is an incredible production-ready tool.
GeoPolars is nowhere near the functionality or stability of GeoPandas, but competition is good and, due to its pure-Rust core, GeoPolars will be much easier to use in WebAssembly.
Note that (for the time being), we started geopolars developement under: https://github.com/pola-rs/geopolars/tree/dev
We were able to do a full port of GFQL from pandas to polars, cypher graph queries on dataframes, including both our CPU + GPU modes, and hit massive speedups: https://www.graphistry.com/blog/cypher-on-polars-cpu-gpu-gra...
It's been impressive!
Yes and no, its not replacing the reason why pandas was popular ie data scientists, but it a full replacement of its pipeline usage, And I would saw also beating out spark
my understanding is Polars is faster, scales better without using external solutions, better API, +Rust. Pandas wins if you want to use what the vast majority of folks are using and have used in the past. Probably has a more complete set of helpers / recipes for the little things you bump into when using it thoroughly, but in the age of LLMs, I think that's minor.
>Pandas wins if you want to use what the vast majority of folks are using
Vast majority of skilled developers are now using Polars, unless they are constrained by lack of Narwhals support in their third-party library of choice (e.g. Great Expectations, SHAP). That's the more important trend to follow.
It’s the API that gave me the push to leave Pandas. 10 or more years of occasional Pandas use and I still had to google for any non-trivial queries.
In that regard, I’m still waiting for a credible jq replacement…
Replacement for jq: https://github.com/01mf02/jaq
Learn SQL and interface with Duck. You will be 100x faster than Pandas/Polaris duo at fraction of memory. Also SQL is supported literally everywhere with a much more capable than Pandas API. Duck outputs to a Pandas Dataframe, but just treat that like a dictionary. Do all your processing, filtering and aggregation in Duck.
Also, try fx.wtf as a replacement for jq. it comes with a in-built tui viewer that supports vi-keybindings. Ecmascript is built into fx.wtf so you can query the JSON with JS notation (where JSON was born). You can use any JS functions including map/reduce/filter or perform any kind of transformation instead of learning jq dsl that you will forget tomorrow.
Avoiding pandas developers is a great reason to use polars imo
It has been for me. I greatly prefer the API, it fits my mental model much better. Give it a try!
See "Pandas should go extinct": https://news.ycombinator.com/item?id=49668198
tl;dr yes
There are awkward things. For example, if you ingest a nanosecond resolution timestamp, there's no way to re-export that out of the Polars dataframe with nanosecond resolution.
A bit surprised about the datafusion results from the post, I have tried it time and time again, but datafusion has always been the leading/trading blowers with polars for our workfloads with duckdb being vastly slower.
There are some benchmarks for the previous version https://benchmark.clickhouse.com/#system=+ti%20rud|Dusa,s|PD...
Doing the benchmarks for 2.0 on the large AWS metal machines at small data sizes (SF=10) really opened my eyes that we have some low-hanging fruit in Polars when it comes to optimizing our constant overhead for smaller queries.
For example our join currently does a full partition into T partitions, for each of the T threads. Overall we create T^2 partitions, which on a 192-core machine is non-trivial. Great if you have a ton of data to feed that with, but if you 'only' have a few dozen million rows it becomes rather small. This is the primary reason we saw in the benchmarks that Polars pinned to 32 threads beats 192 thread Polars at SF=10.
I'll be working on improving that soon. I expect that to have a big impact on SF=10, and a decent impact on ClickBench, which sits between SF=10 and SF=100 in terms of rows.
I use Polars 2.0(rc) to (pre)calculate billions of weather scores on https://therno.com and it has been a lifesaver
Happy that I can upgrade to 2.0 final tonight.
out of core sounds awesome! biggest thing that forced me to switch from pandas/polars to other solutions back in the day.
TIL that Polars supports SQL. Amazing.
At what point can we say that Polars is basically an in-memory database? (Genuine question)
It did since the first releases, but it was limited.
Now it seems they want to go head to head with DuckDB.
Love competition in the local data SQL space, makes everyone better
Although I find pandas a bit aggravating in many ways, for myself and my equally idiotic laboratory scientist pals, seems that it is the default way you might interface with other libraries like SciPy (i.e. they expect things as NumPy arrays or pandas dataframes). Is this a real issue or will most things happily accept a polars dataframe? We don’t work with such large datasets that speed is likely a huge concern tbh.
A coincidence with the fact duckdb is supposed to release 2.0.0 very soon? :)
Actually yes. We had already been planning to do 2.0 for a long time. We originally said we'd move on from 1.x quickly when released 1.0 but ended up staying at 1.x much longer than intended.
From a quick check our first PRs were merged to the 2.0 branch in June:
Are they competing for something?
> first class SQL support, which together with the performance improvements has Polars leading DataFusion and DuckDB in TPC-H and TPC-DS1 benchmarks,
Apparently about something ))
Thing I care about most is whether the old eager-vs-lazy footguns got cleaned up. Half my bugs were a stray collect() in a loop killing the query plan.
Been using polars for over a year now, it is fantastic.
How does Polars relate to DataFusion these days? There's no reason for them not to converge into a single ecosystem, is there?
It does not use datafusion. https://github.com/pola-rs/polars/issues/6197
My team is all moving over to polars and DuckDB
What was your previous setup / stack?
finally long time coming
cant wait to upgrade my Quant trading bot
I'm also using it for BlockRotate, it's time to upgrade.
I can use pandas to clean a dataset, but each cleaning task is usually one line of code. OTOH, With DuckDB with one SQL statement I can replace 40+ lines of polars/pandas. You may reply, SQL isn't as easy to understand! Fair point, it's a declarative language... which is why I use Malloy. Malloy is to TypeScript as Javascript is to SQL. Malloy is much easier to read and write (just as TypeScript is) because it has a built in semantic model -- all the joins, measures, and dimensions are done in one place.
Here is an example [1] of visualizing college football games. Here are all the queries, and semantic model that power all the visualizations [2] Here is the AI generated typescript/react that does the visualizations [3]. The Malloy ecosystem has Malloyyo and Publisher which are replacements for PowerBI and Tableau and Looker. Here is another example for visualizing global trade [4].
[1] - https://mrtimo.github.io/cfb-games/games-2026.html?week=Week... [2] - https://github.com/mrtimo/cfb-games/blob/main/drives.malloy [3] - https://github.com/mrtimo/cfb-games/blob/main/dashboards/gam... [4] - https://tradeexplorer.org/