NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #419 most downloaded on PyPI
Blazingly fast DataFrame library
Last release today
20 Sep 2026
Ships fairly regularly
a new release about every 4 weeks
Nearly every release is documented
notes for 59 of the last 60 stable releases
67 versions withdrawn
withdrawn after publishing
6 years old
434 releases · first in 2021
One column per quarter.
Use BitmapBuilder in yet more places
linear_space (#20678)read_excel and read_ods (#20845)unnest (#20846)Decimal dtype (#20802)SurrealDB and AsyncSurrealDB classes) in read_database (#20799)nest-asyncio in favor of custom logic (#20793)lakefs:// URI for delta scanner (#20757)numpy.float16 values (as Float32) (#20769)column_widths override autofit in write_excel (#20893)log and exp for Decimal type (#20888)join_where or cross join + filter (#20865)Decimal value for fill_null(strategy="one") (#20844)assert that panics on group_by followed by head(n), where n is larger then the frame height (#20819)+ between themselves (#20825)InvalidHeaderValue scanning from S3 on Windows (#20820)clip for Decimal returning wrong values (#20814)str functions (#20781)Decimal dtype (#20802)is_in values to be given as custom Collection (#20801)pl.repeat_by() (#20787)POLARS_VERBOSE (#20797)fallocate() is not permitted (#20796)assert_series_equal for infinities (#20763)legislators-historical.csv (#20858)AExpr::Gather (#20862)ContainsMany to ContainsAny (#20785)Thank you to all our contributors for making this release possible! @alexander-beedie, @arnabanimesh, @braaannigan, @burakemir, @coastalwhite, @etiennebacher, @ion-elgreco, @itamarst, @lukemanley, @mcrumiller, @nameexhaustion, @orlp, @ritchie46 and @stinodego
Update "group\_by\_rolling" (deprecated) to "rolling" in user guide
str.to_decimal keyword-only (#20570)concat_arr (#20681)<dyn SeriesTrait as AsRef<ChunkedArray<T>> (#20664)verify_dict_indices SIMD (#20623)zlib-rs by default and use zstd::with_buffer (#20614)NORMALIZE string function (#20705)GroupsProxy/GroupsPosition to be sliceable and cheaply cloneable (#20673)str.normalize() (#20483)write_excel (#20638)DuplicateError if given a pyarrow Table object with duplicate column names (#20624)index_of() function to Series and Expr (#19894)sqlparser-rs, enabling "LEFT" keyword to be optional for anti/semi joins in SQL queries (#20576)cat.starts_with/cat.ends_with (#20257)allow_invalid_certificates being ignored in storage_options (#20744)map_groups returning all-NULL column (#20743)unique(maintain_order=True) raising InvalidOperationError for null array (#20737)Series.n_unique raising for list of struct (#20724)head() returning extra rows (#20722)schema_overrides in read_csv (#20672)map_elements ignoring skip_nulls=True for struct dtype (#20668)to_arrow() on filtered unit height DataFrame (#20656).default to azure credential provider scope URL (#20651)join_asof panicking for invalid tolerance input (#20643)Int128 dtype serialization (#20629)read_excel and read_ods support reading from raw bytes for all engines (#20636)LIKE and ILIKE operators support multi-line matches (#20613)with_columns() on unit height LazyFrame (#20584)tenant_id to CredentialProviderAzure if given (#20583)str.strip_chars_* methods (#20565)read_excel "engine_options" and "read_options" docstring (#20661)Field (#20625)DataFrame join examples (#20587)NEW_MULTIFILE (#20730)fastexcel release (#20703)MultiScanExec node (#20648)LazyFrame.fill_null (#20558)Thank you to all our contributors for making this release possible! @Jesse-Bakker, @MarcoGorelli, @MoizesCBF, @SamuelAllain, @alexander-beedie, @bschoenmaeckers, @coastalwhite, @eitsupi, @etiennebacher, @itamarst, @jqnatividad, @lukemanley, @mcrumiller, @nameexhaustion, @orlp, @ritchie46 and @stinodego
Collapse expanded filters in eager
IR::DataFrame (#20492)arg_sort fast path onto Column (#20437)Int128 IO support for csv & ipc (#20535)cs.by_dtype and col (#20491)read_excel and read_ods (#20476)Iterable[bool] in Series.filter (#20431)sum_horizontal with boolean inputs (#20531)unique fast path for empty categoricals (#20536)Int128 operations (#20515)ignore_nulls is respected in horizontal sum/mean (#20469)Int128 testing and related fixes (#20494)unique() for empty DataFrames (#20411)list.min and list.max for list[i128] (#20488)align_frames with single row panicking (#20466)sum_horizontal with boolean dtype (#20459)add_business_days, millennium, century and combine methods in Series.dt namespace (#20436)DataFrame.cast (#20532)join_asof (#20509)Expr.all description of Kleene logic (#20409)floor/ceil on integers (#20479)TypeCheckRule to the optimizer (#20425)Thank you to all our contributors for making this release possible! @Biswas-N, @IndexSeek, @Prathamesh-Ghatole, @Terrigible, @alexander-beedie, @brifitz, @coastalwhite, @dependabot, @dependabot[bot], @jqnatividad, @lukemanley, @mcrumiller, @orlp, @ritchie46 and @siddharth-vi
Order observability optimizations
Int128Type (#20232)BytecodeParser on introspection of incompatible UDFs (#20280)read_excel and read_ods (#20430)dt.replace (#19708)DefaultAzureCredential() (#20384)bin.reinterpret (#20263)Schema (#20267)cat.len_chars and cat.len_bytes (#20211)to_physical_repr of nested types (#20413)mmap crash under Emscripten (#20418)new_columns in scan_csv with compressed file (#20412)Series.dt.add_business_days (#20402)from_repr (#20351)from_physical for List/Array (#20311)SELECT 1 FROM (#20241)--extra-index-url from CUDA 12 packages (#20381)fork warning (#20258)project.dynamic = ["version"] to pyproject.toml (#20345)pyo3 and numpy crates to version 0.23 (#20111)select(len()) (#20343)pl.List and pl.Array by default (#20319)CsvParseOptions (#20302)FunctionCastOptions and conservative IR-level cast type-checking (#20286)Int128Type (#20232)Thank you to all our contributors for making this release possible! @Jesse-Bakker, @Terrigible, @ZemanOndrej, @alexander-beedie, @balbok0, @beckernick, @bschoenmaeckers, @coastalwhite, @georgestagg, @hamdanal, @haocheng6, @kszlim, @lukemanley, @mcrumiller, @nameexhaustion, @noexecstack, @orlp, @ptiza, @r-brink, @ritchie46, @rodrigogiraoserrao, @stijnherfst, @stinodego, @tswast and @zero-stroke
Fix incorrect lazy select(len()) with some select orderings
select(len()) with some select orderings (#20222)scratch.is_empty() (#20219)Thank you to all our contributors for making this release possible! @nameexhaustion and @ritchie46
Deprecate ddof parameter for correlation coefficient
Series construction from subclasses of standard Python types (#20166)Series for bytes/binary data 10x faster when dtype not explicitly set (#20157)DataFrame cols constructed from Python Enum values (#20180)maintain_order parameter to joins (#20026)to_datetime / strftime to automatically parse dates with single-digit hour/minute/second (#20144)to_struct() without a list of field names (#20158)pl.select (#20091)write_delta (#20092)bitwise_xor for ScalarColumn (#20140)u64 (#20167)is_elementwise_top_level (#20177)arg_sort for Null series (#20135)*_horizontal functions (#20130).get_column after drop_in_place (#20120)is_in() in when()/then() for full-streaming (#20052)by param description for rolling_*_by functions (#19715)sqlparser-rs from version 0.49 to 0.52 (#20110)memmap2 to version 0.9 (#20105)object_store to version 0.11 (#20102)fs4 to version 0.12 (#20101)thiserror to version 2 (#20097)atoi_simd to version 0.16 (#20098)chrono-tz to 0.10 (#20094)ndarray to 0.16 (#20093)nightly-2024-11-28 (#20064)Thank you to all our contributors for making this release possible! @DzenanJupic, @MarcoGorelli, @YichiZhang0613, @alexander-beedie, @coastalwhite, @dependabot, @dependabot[bot], @flowlight0, @henryharbeck, @iharthi, @ion-elgreco, @jqnatividad, @lukapeschke, @lukemanley, @mcrumiller, @nameexhaustion, @ptiza, @ritchie46, @siddharth-vi, @stijnherfst, @stinodego and @wsyxbcl
Remove note about guaranteed left join order
Config instances (#20053)Enum init (#20060)Enum dtype init from standard Python enums (#19997)drop_nans method to DataFrame and LazyFrame (#20029)hist binning around breakpoints (#20054)Series.hist with bin_count when all values are the same (#20034)hist panicking on out of bounds index (#20016)list.len() for masked-out rows (#19999)collect_schema() for fill_null() after an aggregation expression in group-by context (#19993)row_by_key typing (#19888)Thank you to all our contributors for making this release possible! @alexander-beedie, @coastalwhite, @gab23r, @lukemanley, @mcrumiller, @nameexhaustion, @ritchie46, @siddharth-vi, @stijnherfst and @stinodego
Reduce the size of row encoding UTF-8
pl.List (#19907)to_string for non-Duration dtypes and raise an informative error (#19977)pl.concat_arr to concatenate columns into an Array column (#19881)dt.to_string (#19840)gather_every for Scalar (#19964)numpy arrays (#19895)ShapeError error message on dataframe creation (#19901)to_torch) and Jax Arrays (to_jax) (#19862)scan_parquet().with_row_index() with hive partitioning enabled (#19865)LazyFrame.explode() (#19860)rows_by_key returning key tuples with elements in wrong order (#19486)List element truncation ellipses respect ASCII* table formats (#19835)Series.bottom_k docstring (#19947)clip() (#19875) ) tweak for HTML rendering (#19864)with_columns test (#19844)Thank you to all our contributors for making this release possible! @MarcoGorelli, @alexander-beedie, @barak1412, @coastalwhite, @etiennebacher, @ion-elgreco, @itamarst, @lukemanley, @mcrumiller, @mhogervo, @nameexhaustion, @orlp, @ritchie46, @stijnherfst and @stinodego
Increase default async thread count for low core count systems
to_string, ergonomic/perf improvement, tz-aware Datetime bugfix (#19697)to_string, ergonomic/perf improvement, tz-aware Datetime bugfix (#19697)is_literal method to expression meta namespace (#19773)read_database(…,iter_batches=True) type annotations (#19832)row_by_keys in the to_dict documentation (#19767)Column (#19736)Thank you to all our contributors for making this release possible! @MarcoGorelli, @TNieuwdorp, @YichiZhang0613, @alexander-beedie, @braaannigan, @coastalwhite, @engylemure, @gab23r, @iliya-malecki, @ion-elgreco, @itamarst, @jackxxu, @nameexhaustion, @orlp, @ritchie46, @rodrigogiraoserrao and @sn0rkmaiden
Add IPC source node for new streaming engine
selector & col expansion (#19742)meta.is_column to API docs (#19744)Thank you to all our contributors for making this release possible! @alexander-beedie, @coastalwhite, @etiennebacher, @itamarst, @nameexhaustion, @orlp, @ritchie46 and @rodrigogiraoserrao
Improve DataFrame.sort().limit/top_k performance
DataFrame.sort().limit/top_k performance (#19731)read_database (#19733)n_chunks typing (#19727)removeprefix, removesuffix, and zfill in map_elements (#19672)replace in map_elements (#19668)RIGHT JOIN, fix an issue with wildcard aliasing (#19626)& with IEJoin (join_where) (#19552)is_between range predicate with IEJoin operations (join_where) (#19547)cls for to_python (#19726)ELSE clause should be implicitly NULL when omitted (#19714)n_chunks typing (#19727)NoDataError raised consistently between engines for Excel reads (#19712)list.to_struct to be elementwise when width is fixed (#19688)cast (#19657)explode() in agg() (#19629)mean_horizontal raises on non-numeric input (#19648).struct.with_fields inside list.eval (#19617)scan_parquet().with_row_index() with non-zero slice or with streaming collect (#19609)credential_provider is given (#19589)explode operation in the array namespace (#19163)replace and replace_all docstring explanation of the "$" character with reference to capture groups (vs use as a literal) (#19529)nightly-2024-10-28 (#19492)Column for the {try,}_apply_columns{_par,} functions on DataFrame (#19683)@scalar-opt (#19666)std::ops::Bit... (#19673)Column into polars-expr (#19660)explode operation in the array namespace (#19163)Column::Partitioned variant (#19557)test_rolling_by_integer not using parameterized dtype (#19555)mindebug-dev rust profile (#19524)Thank you to all our contributors for making this release possible! @3tilley, @HansBambel, @MarcoGorelli, @alexander-beedie, @barak1412, @braaannigan, @cmdlineluser, @coastalwhite, @corwinjoy, @dependabot, @dependabot[bot], @eitsupi, @janpipek, @jqnatividad, @letkemann, @max-muoto, @nameexhaustion, @orlp, @ritchie46, @rodrigogiraoserrao, @siddharth-vi, @stinodego and @wence-
Updates error message in csv parser to recommend schema\_overrides instead of deprecated dtypes argument
dt.add_business_days keyword-only (#19428)expand_columns (#19469)read_database typing (#19444)include_index for pandas series (#19453)credential_provider argument to more read functions (#19421)scan_iceberg (#19388)to_physical (#19474)ColumnNotFound error (#19473)list.to_struct when fields are passed (#19439)expand_columns (#19469)eq/ne_missing also compares outer validity (#19443)escape_regex example (#19440)ColumnNotFound when using pl.element() inside list.eval (#19438).join(..., how="left").head(N) if N <= left_df.height() and there are duplicate matches (#19422)ASCII* table formats do not use the UTF8 ellipsis char when truncating rows/cols/values (#19404)datetime (#19459)read_json (#19425)examples folder in favor of the user guide (#19430)Thank you to all our contributors for making this release possible! @alexander-beedie, @cmdlineluser, @coastalwhite, @corleyma, @corwinjoy, @dvillaveces, @eitsupi, @gab23r, @janscholten, @nameexhaustion, @orlp, @ritchie46, @siddharth-vi, @stinodego and @wakabame
Improve var/cov/corr performance
Schema improvements (equality/init dtype checks) (#19379)escape_regex operation to the str namespace and as a global function (#19257)read_database_uri typing (#19334)include_file_paths and with_row_index for streaming CSV scan (#19394)gather_with_series() and friend (#19383)str.json_decode (#19347)is_{min,max}_value_exact when set to true (#19344)include_file_paths (#19341)0.21 (#19376)AlignedBytes types (#19308)Thank you to all our contributors for making this release possible! @alexander-beedie, @barak1412, @benrutter, @coastalwhite, @corwinjoy, @itamarst, @max-muoto, @nameexhaustion, @orlp, @ritchie46, @stinodego, @wence- and @wolfgang-noichl
Add/fix unordered row decode, change unordered format
18% / 25% on 1.9.0 (#19124)~17% (#19088)bit_count and bitwise &, |, and xor operators (#19114)credential_provider argument for scan_parquet (#19271)SortingColumns for ints (#19251)read_ods (#19202)read_excel (#18253)HAVING outside of GROUP BY should raise a suitable SQLSyntaxError (#19320)from_dicts typing/signature (#19322)unpivot (#19313)as_struct (#19280)read_database takes advantage of Arrow return from a duckdb_engine connection when using a SQLAlchemy Selectable (#19255)lit(_) != (#19246)Series from arrow struct (#19218)read_excel (#18253)write_excel (#19029)IO[bytes] instance (#19154)(eq|ne)_missing on List/Array types (#19155)~17% (#19088)DatetimeOwned to ChunkedArray (#19094)schema arg is propagated to IR (#19084)QUANTILE_CONT and QUANTILE_DISC functions (#19272)as_struct (#19116)Series.first,last,approx_n_unique to docs (#19146)parquet-format-safe to polars-parquet-format (#19275)get_list_builder infallible (#19217)pl.repeat part of the IR (#19152)Thank you to all our contributors for making this release possible! @Bidek56, @MarcoGorelli, @Rashik-raj, @adamreeve, @alexander-beedie, @alonme, @balbok0, @coastalwhite, @deanm0000, @dependabot, @dependabot[bot], @eitsupi, @etrotta, @itamarst, @jbutterwick, @joelostblom, @kenkoooo, @khalidmammadov, @laurentS, @mcrumiller, @mscolnick, @nameexhaustion, @orlp, @pomo-mondreganto, @ritchie46, @rodrigogiraoserrao, @siddharth-vi, @stinodego, @sunadase and @wence-
Bitwise operations / aggregations
insert_column to take expressions (#19024)strict param to eager/lazy frame "rename" (#19017)schema arg in read/scan_parquet() (#19013)include_file_paths parameter to read_parquet (#19008)allow_missing_columns option to read/scan_parquet (#18922)must_flush flag is not reset (#19046)missing equality (#19031)ne_missing and eq_missing operations for struct instead of null (#18930)when().then().otherwise() on struct when both result are broadcast (#19000)when().then().else() on structs when using first()\last() (#18969)lit().shrink_dtype() broadcasting (#18958)cumulative_eval (#18959)into_py (#18960)Expr.over with order_by did not take effect if group keys were sorted (#18947)with_row_index() to previously collected lazy scan does not take effect (#18913)is_not_nan description (#18985)nightly-2024-09-29 (#19006)simd-json to 0.14 (#18999)schema arg in read/scan_parquet as unstable (#19018)test_lazy_parquet::test_row_index (#19019)allow_missing_columns in error message when column not found (parquet) (#18972)ChunkCompare into Eq and Ineq variants (#18963)Thank you to all our contributors for making this release possible! @LukasFolwarczny, @Plutone11011, @aleexharris, @alexander-beedie, @barak1412, @coastalwhite, @dependabot, @dependabot[bot], @edwinvehmaanpera, @kgv, @mcrumiller, @nameexhaustion, @orlp, @ritchie46, @rodrigogiraoserrao, @stinodego and @xhiroga
Improve rename performace for Lazy API
Thank you to all our contributors for making this release possible! @coastalwhite, @npielawski, @orlp, @ritchie46 and @siddharth-vi
Properly calculate duration units
with_column_unchecked take Column (#18863)Thank you to all our contributors for making this release possible! @MarcoGorelli, @coastalwhite, @mcrumiller, @ritchie46 and @rodrigogiraoserrao
Support arithmetic between Series with dtype list
ruff lint rule sets (#18721)x=alt.X(a, axis=alt.Axis(labelAngle=30))) (#18836)select() was done (#18843)write_csv (#18845)join argument checks (#18847)arr.to_struct (#18804)streaming=True (#18766)cum_max using exception text of cum_min for invalid dtype (#18780)schema in read_csv function (#18759)lit docstrings (#18756)docs directory hierarchy (#18773)over docs, add example with order_by (#18796)polars-python crate (#18835)NodeTraverser struct public (#18822)Column instead of Series (#18664)ruff lint rule sets (#18721)Thank you to all our contributors for making this release possible! @3ok, @Manishearth, @MarcoGorelli, @adamreeve, @alexander-beedie, @barak1412, @beckernick, @bradfordlynch, @coastalwhite, @deanm0000, @eitsupi, @i64, @itamarst, @mcrumiller, @nameexhaustion, @orlp, @r-brink, @ritchie46, @rodrigogiraoserrao, @squnit, @stinodego and @t-ded
Revert automatically turning on Parquet prefiltered
Thank you to all our contributors for making this release possible! @ankane, @attila-lin, @coastalwhite, @eitsupi, @nameexhaustion, @orlp and @ritchie46
Add support for IO[bytes] and bytes in scan_{...} functions
IO[bytes] and bytes in scan_{...} functions (#18532)ColumnChunkMetadata (#18615)ColumnChunkMetadata (#18584)parallel=prefiltered for auto (#18514)PlSmallStr impl from Arc<str> to compact_str (#18508)is_null().all() and similar expressions to use null_count() (#18359)BytecodeParser for upcoming Python 3.13 (#18677)IO[bytes] and bytes in scan_{...} functions (#18532)DataFrame.write_parquet() (#18652)replace/replace_strict is not a mapping (#18492)list.eval in certain cases (#18570)map_elements for List return dtypes (#18567)read_database cursor result, raising DuplicateError if found (#18548)maintain_order=True (#18561)map_elements for List types (#18542)align_frames result when the alignment column contains NULL values (#18521)write_database not passing down "engine_options" when using ADBC (#18451)or and xor operations (#18512)assert_frame_not_equal and assert_series_not_equal raise on mismatched input types (#18402)Worksheet definition in write_excel type annotations (#18452)testing.assert_* functions (#18494)streaming argument in test_parquet_slice_pushdown_non_zero_offset (#18529)PlSmallStr impl from Arc<str> to compact_str (#18508)Thank you to all our contributors for making this release possible! @0xbe7a, @MarcoGorelli, @WbaN314, @adamreeve, @alexander-beedie, @alonme, @barak1412, @coastalwhite, @dependabot, @dependabot[bot], @eitsupi, @henryharbeck, @ion-elgreco, @krasnobaev, @megaserg, @nameexhaustion, @ohanf, @orlp, @philss, @r-brink, @ritchie46, @skellys, @squnit, @stinodego, @wence- and @yarimiz
Deprecate serialize json for LazyFrame
These API's were marked unstable and are allowed to change.
DELTA_LENGTH_BYTE_ARRAY decoding (#18299)time/timedelta literals (#18223)~40% (#18197)str.replace_many (#18214)read_database (#18277)explode as gather (#18431)scan_parquet(parallel='prefiltered') problems (#18278)upsample only have to be sorted within groups (#18264)hist when bin_count specified (#16942)SQL set op syntax (#18205)include_file_paths (#18255)eager=True (#18379)group_by_dynamic (#18415)Series methods to API reference (#18312)DataFrame.__getitem__ and Series.__getitem__ (#18309)coalesce behaviour in join_asof (#18273)Expr.shuffle differentiating from df method (#18266)bin.size expr docstring (#18222)DataFrame.map_rows (#18227)nightly-2024-08-26 (#18370)py-polars crate (#18204)test_read_database_cx_credentials (#18220)Thank you to all our contributors for making this release possible! @BartSchuurmans, @ChayimFriedman2, @MarcoGorelli, @StepfenShawn, @agossard, @alexander-beedie, @cgbur, @coastalwhite, @corwinjoy, @deanm0000, @henryharbeck, @ion-elgreco, @jqnatividad, @krasnobaev, @liufeimath, @markxwang, @mcrumiller, @nameexhaustion, @orlp, @ritchie46, @stinodego, @sunadase, @thomascamminady and @wence-
Improve binview extend/ifthenelse
Arc<Vec<_>> instead of Arc<[_]> for paths and hive partitions (#18066)FixedSizeBinary (#18059)FixedSizeBinary to BinaryView cast (#18043)read_excel and read_ods (#18078)Config state (#18151)filter=None (#18139)include_index=False (the default) (#18133)read_csv (#18131)to_titlecase was too narrowly defined (#18122)write_excel column totals, don't forget to include any row-total cols (#18042)select(len()) for compressed files (#18067)sink_ipc_cloud panicking with runtime error (#18091)CloudWriter to use buffer before making requests (#18027)cfg(feature) for shrink_dtype (#18038)lazy docstring (#18178)Thank you to all our contributors for making this release possible! @EricTulowetzke, @KDruzhkin, @MarcoGorelli, @Vincenthays, @alexander-beedie, @coastalwhite, @davanstrien, @deanm0000, @ember91, @kylebarron, @mcrumiller, @nameexhaustion, @orlp, @philss, @ritchie46 and @rosstitmarsh
Integer fast path Parquet dict encoding
Worksheet objects to the write_excel method (#18031)comfy-table version (#18028)Thank you to all our contributors for making this release possible! @alexander-beedie, @coastalwhite, @deanm0000, @nameexhaustion and @ritchie46
Push down slice with non-zero offset to Parquet
.dt.weekday 20x faster (#17992)MemSliceInner enum (#17991)MemSlice (#17983)size method to Expr and Series "bin" namespace (#17924)SQL interface support for PostgreSQL dollar-quoted string literals (#17940)strict argument (#17990)sort method (#17947)COUNT(DISTINCT x) should not include NULL values (#17930)None in pycapsule interface export (#17922)last is never ambiguous with max (#17962)str.contains_any and str.replace_many (#17961)allow_null as replacement (#17969)Thank you to all our contributors for making this release possible! @JamesCE2001, @MarcoGorelli, @alexander-beedie, @coastalwhite, @deanm0000, @deepyaman, @dependabot, @dependabot[bot], @henryharbeck, @kylebarron, @nameexhaustion, @ritchie46 and @wangxiaoying
Better deprecate message for \_import\_from\_c
MemReader to file buffer in Parquet reader (#17712)apply_into_string_amortized instead of apply_to_buffer (#17903)SQL "INTERSECT" and "EXCEPT" set ops (#17835)write_excel (#17757)to_string for Date dtype (#17670)is_in operation on decimal type (#17832)read_excel when using "calamine" engine with the latest fastexcel (#17735)hf:// in read_(csv|ipc|ndjson) functions (#17785)collect_schema (#17761)hf:// (#17682)get_column_index (#17868)glob=False for cloud reads (#17860)write_excel int/float format when using a dark "table_style" (#17869)from_arrow for struct type (#17839)write_excel (#17846)NullArray in Parquet (#17807)named_expr and schema in pl.struct (#17768)write_ipc (#17752)join types for clarity (#17843)read_* functions in Hugging Face section in user guide (#17799)Expr.map_batches (#17789)nightly-2024-07-26 (#17891)uv pip install to verbose (#17901)typos command in make pre-commit for py-polars folder (#17897)typos configuration features (#17800)setuptools (#17726)Thank you to all our contributors for making this release possible! @MarcoGorelli, @Object905, @SandroCasagrande, @alexander-beedie, @atigbadr, @coastalwhite, @deanm0000, @delsner, @dependabot, @dependabot[bot], @henryharbeck, @implicit-apparatus, @jparag, @knl, @kylebarron, @lukapeschke, @mcrumiller, @nameexhaustion, @orlp, @ritchie46, @ruihe774, @stinodego, @szepeviktor and @wence-
Specify tune-cpu \& add more features
__all__ (#17494)sink methods (#17698)read_database issue with batched reads from Snowflake (#17688)value_counts methods based on normalize parameter (#17685)setuptools to fix failing CI (#17695)sink methods (#17698)setuptools to fix failing CI (#17695)ComputeNode in new streaming engine (#17389)Thank you to all our contributors for making this release possible! @5j9, @ByteNybbler, @MarcoGorelli, @alexander-beedie, @coastalwhite, @diegoglozano, @eitsupi, @nameexhaustion, @orlp, @ragyabraham, @ritchie46 and @ruihe774
Fixup "deprecated" directive for DataFrame.melt and LazyFrame.melt
scan functions (#17616)ArrayChunks to optimize codegen of BatchDecoder (#17632)infer_schema parameter to read_csv / scan_csv (#17617)returns_scalar to map_elements (#17613)describe on decimal (#15092)write_database (#17470)pivot_schema (#17611)sort_by_exprs() (#17606)O_CLOEXEC on duplicated file descriptor (#17537)collect in file scan methods (#17532)retries parameter in scan functions not taking effect when it was set to 0 (#17564).list.(get|gather) (#17511)scan_ipc does not go through fsspec (#17495)sink_csv (#17476)plot docs to refer to docstrings (#17504)str.lengths to str.len_bytes in description text (#11577) (#17626)polars.Expr.bin.decode (#17508)read_database_uri docstring (#17536)DataFrame.melt and LazyFrame.melt (#17530)write_parquet_partitioned (#17488)ArrayChunks to optimize codegen of BatchDecoder (#17632)utils to path_utils in polars-io (#17635)with_column method of PyLazyFrame (#17607)style accessor to DataFrame (#17502)is_supported_cloud util (#17493)Thank you to all our contributors for making this release possible! @Julian-J-S, @MarcoGorelli, @alexander-beedie, @anergictcell, @arnabanimesh, @brandon-b-miller, @cmdlineluser, @coastalwhite, @deanm0000, @eitsupi, @flisky, @henryharbeck, @itamarst, @jonaylor89, @moritzwilksch, @nameexhaustion, @orlp, @phi-friday, @r-brink, @rcorty, @ritchie46, @ruihe774, @stinodego, @tylerriccio33 and @wence-
Keep more parallelism when CSE plan cache hits
scan_ipc (#17434)Series.__getitem__ (#17408)read_excel engines (#17448)from_pandas for string columns with missing values (#17397)SQL interface (#17400)slice length no longer allowing None (#17372)SchemaError exception message (#17350)partition_by docstring to match new behavior (#17394)GroupBy.__iter__ docstring to match new behavior (#17383)np.trapz in tests to prepare for NumPy 2.0 (#17387)sink_csv test (#17386)Thank you to all our contributors for making this release possible! @alexander-beedie, @brunobbaraujo, @cmdlineluser, @coastalwhite, @dependabot, @dependabot[bot], @nameexhaustion, @orlp, @phi-friday, @ritchie46, @ruihe774, @sherlockbeard, @stinodego, @tylerriccio33 and @wence-
This is the first major release for Python Polars. Please check out the upgrade guide for help navigating the breaking changes when upgrading to this…
This is the first major release for Python Polars. Please check out the upgrade guide for help navigating the breaking changes when upgrading to this version.
read_excel to "calamine" (#17263)pyproject.toml (#17168)read/scan_parquet to disable Hive partitioning by default for file inputs (#17106)replace functionality into two separate methods (#16921)compression argument as keyword-only (#17084)ModuleUpgradeRequired and PolarsPanicError error, remove InvalidAssert error (#17033)strict parameter in Series constructor (#16939)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)selector XOR set operation, guarantee consistent selector column-order (#16833)infer_schema_length as keyword-only argument in str.json_decode (#16835)set_sorted to only accept a single column (#16800)Series.cut/qcut and update struct field names (#16741)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)LazyFrame.fetch (#17278)size parameter in parametric testing strategies in favor of min_size/max_size (#17128)replace functionality into two separate methods (#16921)DataFrame.melt to unpivot and make parameters consistent with pivot (#17095)dt.mean/dt.median in favor of mean/median (#16888)LazyFrame.with_context in favor of horizontal concatenation (#16860)descending to reverse in top_k methods (#16817)str.concat to str.join and update default delimiter (#16790)arctan2d in favor of arctan2(...).degrees() (#16786)group_by `iteration (#17302)unique performance by adding RangedUniqueKernel for primitive arrays (#17166)unique performance by creating UniqueKernel and improve bool implementation (#17160)compression argument as keyword-only (#17084)if-then-else view kernel (#16993)AND filter into multiple nodes (#16992)arg_sort of row-encoding (#16894)rle_id iteration performance and set sorted flags (#16893)sort for String and Binary types (#16871)split_at in split (#16865)split_at instead of double slice in chunk splits. (#16856)align_ if arrays are aligned (#16850)arg_sort (#16808)dt.offset_by 2x for constant durations (#16728)join if non-coalesced key isn't projected (#16677)dt.truncate 1.5x faster when every is just a single duration (and not an expression) (#16666)NATURAL joins and the COLUMNS function (#17295)str.extract_many expression (#17304)read_excel to "calamine" (#17263)LazyFrame.fetch (#17278)SQL Struct/JSON field access operators (#17226)ORDER BY ALL syntax (#17212)^@ ("starts with"), and ~~,~~*,!~~,!~~* ("like", "ilike") string-matching operators (#17251)SELECT * ILIKE wildcard syntax (#17169)SQL temporal functions STRFTIME and STRPTIME, and typed literal syntax (#17245)round/ceil/floor on integer types (#17241)write_csv/write_json (#14209)get_column DataFrame method (#17176)float_scientific option to write_csv/sink_csv (#17111)Struct field selection in the SQL engine, RENAME and REPLACE select wildcard options (#17109)DataFrame.pivot to allow index=None when values is set (#17126)read/scan_parquet to disable Hive partitioning by default for file inputs (#17106)replace functionality into two separate methods (#16921)DataFrame.melt to unpivot and make parameters consistent with pivot (#17095)explain and show_graph (#17074)pl.col autocompletion for iPython (#17080)read_ndjson (#17068)strict parameter to DataFrame/LazyFrame.drop and fix behavior to default to True (#17044)ModuleUpgradeRequired and PolarsPanicError error, remove InvalidAssert error (#17033)rechunk parameter to read_delta (#16991)json_normalize (#17015)AND filter into multiple nodes (#16992)strict parameter in Series constructor (#16939)INTERSECT and EXCEPT ops (#16960)PerformanceWarning to LazyFrame properties (#16964)collect_schema method to LazyFrame and DataFrame (#16929)lit (#16950)DataFrame.style namespace (#16809)Schema class (#16873)value_counts (#16917)read_csv SQL table reading function defaults (better handle dates) (#16866)VALUES clause and inline renaming of columns in CTE & derived table definitions (#16851)Enum values in lit (#16858).str.to_datetime when values are offset-aware (#16742)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)SQL "SELECT" with no tables, optimise registration of globals (#16836)selector XOR set operation, guarantee consistent selector column-order (#16833)EXTRACT and DATE_PART SQL part abbreviations (#16767)set_sorted to only accept a single column (#16800)group_by iteration and partition_by to always return tuple keys (#16793)read_database_uri passthrough from read_database (#16783)pyxlsb engine from read_excel (#16784)check_order parameter to assert_series_equal (#16778)scan_csv (#16674)INTERVAL handling and improve related error messages, update sqlparser-rs lib (#16744)ORDER BY clause (#16745)pandas and pyarrow objects (#16746)Series.cut/qcut and update struct field names (#16741)date_range to no longer produce datetime ranges (#16734)min_periods as keyword-only for rolling methods (#16738)top_k parameters nulls_last, maintain_order, and multithreaded (#16599)NULLS FIRST/LAST ordering (#16711)INTERVAL strings (#16732)offset arg in truncate and round (#16655)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)str.to_datetime (#16634)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)SQLInterface and SQLSyntax errors (#16635)DIV function support to the SQL interface (#16678)sink_csv fails (#17313)adbc connections in write_database (#17298)list.get for column index (#17276)list.get (#17262)nulls_last parameter in aggregate sort_by (#17249)DataFrame.top_k not handling nulls correctly (#17239)lit to address spurious test failure (#17187)ChainedWhen should not inherit Expr (#17142)fold in certain situations (#17114)Series dunder method type signatures (#17053)sqlalchemy libraries (#17029)sort_by of unequal length (#17026)FAST_EXPLODE_LIST metadata (#16951)extend() (#16890)should_rechunk check (#16852)read_excel and read_ods return identical frames across all engines when given empty spreadsheet tables (#16802)read_excel (#16840)top_k/bottom_k and fix a variety of bugs (#16804)DATE_PART SQL syntax/parsing, improve some error messages (#16761)pl. qualifier for inner dtypes in to_init_repr (#16235)assert_series_equal when categorical_as_str=True (#16700)read_database check for SQLAlchemy async Session objects (#16680)selector set ops (#17299)CAST and TRY_CAST functions (#17214)plot namespace as unstable (#17205)concat_list (#17127)DataFrame.unique docstring (#17119)InProcessQuery in docs, mark as unstable (#17097)write_parquet docstring (#16909)select and with_columns to idiomatic form (#16801)DataFrame.limit (#16753)include_nulls in DataFrame.update docstring (#16701)DataFrame.rolling (#16600)Expr/Series.map_elements (#16079)polars.sql docs entry and small docstring update (#16656)pyproject.toml (#17168)< 2.0.0 for now (#17060)iter in list.get (#17286)type_aliases module to _typing (#17282)cargo.toml (#17145)concat_list (#17120)orient="row" in DataFrame constructor when applicable (#16977)Arc from FileCacheEntry (#16870)infer_schema_length as keyword-only argument in str.json_decode (#16835)ChunkedArray::from_chunks_and_dtype (#16697)1.0.0 release (#16705)Thank you to all our contributors for making this release possible! @IvanIsCoding, @JamesCE2001, @JulianCologne, @KDruzhkin, @Kylea650, @MarcoGorelli, @Mottl, @Object905, @SeanTater, @adamreeve, @alexander-beedie, @bertiewooster, @borchero, @c-peters, @coastalwhite, @datapythonista, @datenzauberai, @dependabot, @dependabot[bot], @eitsupi, @flisky, @henryharbeck, @itamarst, @jqnatividad, @lukeshingles, @machow, @marenwestermann, @mcrumiller, @montanarograziano, @nameexhaustion, @orlp, @p3i0t, @ritchie46, @sherlockbeard, @stinodego, @tkellogg, @universalmind303 and @wence-
Remove deprecated parameters in Series.cut/qcut and update struct field names
hive_partitioning parameter default to None, which is automatically enabled for single directory inputs, and disabled otherwise (#17106)replace functionality into two separate functions (#16921)strict parameter to DataFrame/LazyFrame.drop and fix behavior to default to True (#17044)ModuleUpgradeRequired and PolarsPanicError error, remove InvalidAssert error (#17033)strict parameter in Series constructor (#16939)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)selector XOR set operation, guarantee consistent selector column-order (#16833)infer_schema_length as keyword-only argument in str.json_decode (#16835)set_sorted to only accept a single column (#16800)Series.cut/qcut and update struct field names (#16741)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)size parameter in parametric testing strategies in favor of min_size/max_size (#17128)replace functionality into two separate functions (#16921)DataFrame.melt to unpivot and make parameters consistent with pivot (#17095)dt.mean/dt.median in favor of mean/median (#16888)LazyFrame.with_context in favor of horizontal concatenation (#16860)descending to reverse in top_k methods (#16817)str.concat to str.join and update default delimiter (#16790)arctan2d in favor of arctan2(...).degrees() (#16786)AND filter into multiple nodes (#16992)split_at in split (#16865)split_at instead of double slice in chunk splits. (#16856)align_ if arrays are aligned (#16850)arg_sort (#16808)dt.offset_by 2x for constant durations (#16728)join if non-coalesced key isn't projected (#16677)dt.truncate 1.5x faster when every is just a single duration (and not an expression) (#16666)float_scientific option to write_csv/sink_csv (#17111)Struct field selection in the SQL engine, RENAME and REPLACE select wildcard options (#17109)DataFrame.pivot to allow index=None when values is set (#17126)hive_partitioning parameter default to None, which is automatically enabled for single directory inputs, and disabled otherwise (#17106)replace functionality into two separate functions (#16921)DataFrame.melt to unpivot and make parameters consistent with pivot (#17095)pl.col autocompletion for iPython (#17080)strict parameter to DataFrame/LazyFrame.drop and fix behavior to default to True (#17044)ModuleUpgradeRequired and PolarsPanicError error, remove InvalidAssert error (#17033)rechunk parameter to read_delta (#16991)json_normalize (#17015)AND filter into multiple nodes (#16992)strict parameter in Series constructor (#16939)INTERSECT and EXCEPT ops (#16960)PerformanceWarning to LazyFrame properties (#16964)collect_schema method to LazyFrame and DataFrame (#16929)lit (#16950)Schema class (#16873)value_counts (#16917)eq/ne for more FixedSizeLists (#16902)read_csv SQL table reading function defaults (better handle dates) (#16866)VALUES clause and inline renaming of columns in CTE & derived table definitions (#16851)Enum values in lit (#16858).str.to_datetime when values are offset-aware (#16742)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)SQL "SELECT" with no tables, optimise registration of globals (#16836)selector XOR set operation, guarantee consistent selector column-order (#16833)EXTRACT and DATE_PART SQL part abbreviations (#16767)set_sorted to only accept a single column (#16800)group_by iteration and partition_by to always return tuple keys (#16793)read_database_uri passthrough from read_database (#16783)pyxlsb engine from read_database (#16784)check_order parameter to assert_series_equal (#16778)scan_csv (#16674)INTERVAL handling and improve related error messages, update sqlparser-rs lib (#16744)ORDER BY clause (#16745)pandas and pyarrow objects (#16746)Series.cut/qcut and update struct field names (#16741)date_range to no longer produce datetime ranges (#16734)min_periods as keyword-only for rolling methods (#16738)top_k parameters nulls_last, maintain_order, and multithreaded (#16599)NULLS FIRST/LAST ordering (#16711)INTERVAL strings (#16732)offset arg in truncate and round (#16655)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)str.to_datetime (#16634)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)SQLInterface and SQLSyntax errors (#16635)DIV function support to the SQL interface (#16678)ChainedWhen should not inherit Expr (#17142)GetOutput::get_field fallible (#17114)Series dunder method type signatures (#17053)sqlalchemy libraries (#17029)FAST_EXPLODE_LIST metadata (#16951)extend() (#16890)should_rechunk check (#16852)read_excel and read_ods return identical frames across all engines when given empty spreadsheet tables (#16802)read_excel (#16840)top_k/bottom_k and fix a variety of bugs (#16804)DATE_PART SQL syntax/parsing, improve some error messages (#16761)pl. qualifier for inner dtypes in to_init_repr (#16235)assert_series_equal when categorical_as_str=True (#16700)read_database check for SQLAlchemy async Session objects (#16680)concat_list (#17127)DataFrame.unique docstring (#17119)InProcessQuery in docs, mark as unstable (#17097)select and with_columns to idiomatic form (#16801)DataFrame.limit (#16753)include_nulls in DataFrame.update docstring (#16701)DataFrame.rolling (#16600)Expr/Series.map_elements (#16079)polars.sql docs entry and small docstring update (#16656)< 2.0.0 for now (#17060)cargo.toml (#17145)concat_list (#17120)orient="row" in DataFrame constructor when applicable (#16977)Arc from FileCacheEntry (#16870)infer_schema_length as keyword-only argument in str.json_decode (#16835)ChunkedArray::from_chunks_and_dtype (#16697)1.0.0 release (#16705)Thank you to all our contributors for making this release possible! @JulianCologne, @KDruzhkin, @Kylea650, @MarcoGorelli, @Mottl, @Object905, @adamreeve, @alexander-beedie, @bertiewooster, @borchero, @c-peters, @coastalwhite, @datapythonista, @datenzauberai, @dependabot, @dependabot[bot], @eitsupi, @henryharbeck, @itamarst, @lukeshingles, @machow, @marenwestermann, @mcrumiller, @montanarograziano, @nameexhaustion, @orlp, @p3i0t, @ritchie46, @sherlockbeard, @stinodego, @tkellogg, @universalmind303 and @wence-
Remove deprecated parameters in Series.cut/qcut and update struct field names
hive_partitioning parameter default to None, which is automatically enabled for single directory inputs, and disabled otherwise (#17106)replace functionality into two separate functions (#16921)strict parameter to DataFrame/LazyFrame.drop and fix behavior to default to True (#17044)ModuleUpgradeRequired and PolarsPanicError error, remove InvalidAssert error (#17033)strict parameter in Series constructor (#16939)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)selector XOR set operation, guarantee consistent selector column-order (#16833)infer_schema_length as keyword-only argument in str.json_decode (#16835)set_sorted to only accept a single column (#16800)Series.cut/qcut and update struct field names (#16741)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)size parameter in parametric testing strategies in favor of min_size/max_size (#17128)replace functionality into two separate functions (#16921)DataFrame.melt to unpivot and make parameters consistent with pivot (#17095)dt.mean/dt.median in favor of mean/median (#16888)LazyFrame.with_context in favor of horizontal concatenation (#16860)descending to reverse in top_k methods (#16817)str.concat to str.join and update default delimiter (#16790)arctan2d in favor of arctan2(...).degrees() (#16786)AND filter into multiple nodes (#16992)split_at in split (#16865)split_at instead of double slice in chunk splits. (#16856)align_ if arrays are aligned (#16850)arg_sort (#16808)dt.offset_by 2x for constant durations (#16728)join if non-coalesced key isn't projected (#16677)dt.truncate 1.5x faster when every is just a single duration (and not an expression) (#16666)DataFrame.pivot to allow index=None when values is set (#17126)hive_partitioning parameter default to None, which is automatically enabled for single directory inputs, and disabled otherwise (#17106)replace functionality into two separate functions (#16921)DataFrame.melt to unpivot and make parameters consistent with pivot (#17095)pl.col autocompletion for iPython (#17080)strict parameter to DataFrame/LazyFrame.drop and fix behavior to default to True (#17044)ModuleUpgradeRequired and PolarsPanicError error, remove InvalidAssert error (#17033)rechunk parameter to read_delta (#16991)json_normalize (#17015)AND filter into multiple nodes (#16992)strict parameter in Series constructor (#16939)INTERSECT and EXCEPT ops (#16960)PerformanceWarning to LazyFrame properties (#16964)collect_schema method to LazyFrame and DataFrame (#16929)lit (#16950)Schema class (#16873)value_counts (#16917)eq/ne for more FixedSizeLists (#16902)read_csv SQL table reading function defaults (better handle dates) (#16866)VALUES clause and inline renaming of columns in CTE & derived table definitions (#16851)Enum values in lit (#16858).str.to_datetime when values are offset-aware (#16742)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)SQL "SELECT" with no tables, optimise registration of globals (#16836)selector XOR set operation, guarantee consistent selector column-order (#16833)EXTRACT and DATE_PART SQL part abbreviations (#16767)set_sorted to only accept a single column (#16800)group_by iteration and partition_by to always return tuple keys (#16793)read_database_uri passthrough from read_database (#16783)pyxlsb engine from read_database (#16784)check_order parameter to assert_series_equal (#16778)scan_csv (#16674)INTERVAL handling and improve related error messages, update sqlparser-rs lib (#16744)ORDER BY clause (#16745)pandas and pyarrow objects (#16746)Series.cut/qcut and update struct field names (#16741)date_range to no longer produce datetime ranges (#16734)min_periods as keyword-only for rolling methods (#16738)top_k parameters nulls_last, maintain_order, and multithreaded (#16599)NULLS FIRST/LAST ordering (#16711)INTERVAL strings (#16732)offset arg in truncate and round (#16655)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)str.to_datetime (#16634)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)SQLInterface and SQLSyntax errors (#16635)DIV function support to the SQL interface (#16678)GetOutput::get_field fallible (#17114)Series dunder method type signatures (#17053)sqlalchemy libraries (#17029)FAST_EXPLODE_LIST metadata (#16951)extend() (#16890)should_rechunk check (#16852)read_excel and read_ods return identical frames across all engines when given empty spreadsheet tables (#16802)read_excel (#16840)top_k/bottom_k and fix a variety of bugs (#16804)DATE_PART SQL syntax/parsing, improve some error messages (#16761)pl. qualifier for inner dtypes in to_init_repr (#16235)assert_series_equal when categorical_as_str=True (#16700)read_database check for SQLAlchemy async Session objects (#16680)concat_list (#17127)DataFrame.unique docstring (#17119)InProcessQuery in docs, mark as unstable (#17097)select and with_columns to idiomatic form (#16801)DataFrame.limit (#16753)include_nulls in DataFrame.update docstring (#16701)DataFrame.rolling (#16600)Expr/Series.map_elements (#16079)polars.sql docs entry and small docstring update (#16656)< 2.0.0 for now (#17060)concat_list (#17120)orient="row" in DataFrame constructor when applicable (#16977)Arc from FileCacheEntry (#16870)infer_schema_length as keyword-only argument in str.json_decode (#16835)ChunkedArray::from_chunks_and_dtype (#16697)1.0.0 release (#16705)Thank you to all our contributors for making this release possible! @JulianCologne, @KDruzhkin, @Kylea650, @MarcoGorelli, @Mottl, @Object905, @alexander-beedie, @bertiewooster, @borchero, @c-peters, @coastalwhite, @datenzauberai, @dependabot, @dependabot[bot], @henryharbeck, @itamarst, @machow, @marenwestermann, @mcrumiller, @montanarograziano, @nameexhaustion, @orlp, @p3i0t, @ritchie46, @sherlockbeard, @stinodego, @tkellogg, @universalmind303 and @wence-
Remove deprecated parameters in Series.cut/qcut and update struct field names
strict parameter in Series constructor (#16939)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)selector XOR set operation, guarantee consistent selector column-order (#16833)infer_schema_length as keyword-only argument in str.json_decode (#16835)set_sorted to only accept a single column (#16800)Series.cut/qcut and update struct field names (#16741)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)dt.mean/dt.median in favor of mean/median (#16888)LazyFrame.with_context in favor of horizontal concatenation (#16860)descending to reverse in top_k methods (#16817)str.concat to str.join and update default delimiter (#16790)arctan2d in favor of arctan2(...).degrees() (#16786)AND filter into multiple nodes (#16992)split_at in split (#16865)split_at instead of double slice in chunk splits. (#16856)align_ if arrays are aligned (#16850)arg_sort (#16808)dt.offset_by 2x for constant durations (#16728)join if non-coalesced key isn't projected (#16677)dt.truncate 1.5x faster when every is just a single duration (and not an expression) (#16666)json_normalize (#17015)AND filter into multiple nodes (#16992)strict parameter in Series constructor (#16939)INTERSECT and EXCEPT ops (#16960)PerformanceWarning to LazyFrame properties (#16964)collect_schema method to LazyFrame and DataFrame (#16929)lit (#16950)Schema class (#16873)value_counts (#16917)eq/ne for more FixedSizeLists (#16902)read_csv SQL table reading function defaults (better handle dates) (#16866)VALUES clause and inline renaming of columns in CTE & derived table definitions (#16851)Enum values in lit (#16858).str.to_datetime when values are offset-aware (#16742)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)SQL "SELECT" with no tables, optimise registration of globals (#16836)selector XOR set operation, guarantee consistent selector column-order (#16833)EXTRACT and DATE_PART SQL part abbreviations (#16767)set_sorted to only accept a single column (#16800)group_by iteration and partition_by to always return tuple keys (#16793)read_database_uri passthrough from read_database (#16783)pyxlsb engine from read_database (#16784)check_order parameter to assert_series_equal (#16778)scan_csv (#16674)INTERVAL handling and improve related error messages, update sqlparser-rs lib (#16744)ORDER BY clause (#16745)pandas and pyarrow objects (#16746)Series.cut/qcut and update struct field names (#16741)date_range to no longer produce datetime ranges (#16734)min_periods as keyword-only for rolling methods (#16738)top_k parameters nulls_last, maintain_order, and multithreaded (#16599)NULLS FIRST/LAST ordering (#16711)INTERVAL strings (#16732)offset arg in truncate and round (#16655)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array type instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)str.to_datetime (#16634)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)SQLInterface and SQLSyntax errors (#16635)DIV function support to the SQL interface (#16678)FAST_EXPLODE_LIST metadata (#16951)extend() (#16890)should_rechunk check (#16852)read_excel and read_ods return identical frames across all engines when given empty spreadsheet tables (#16802)read_excel (#16840)top_k/bottom_k and fix a variety of bugs (#16804)DATE_PART SQL syntax/parsing, improve some error messages (#16761)pl. qualifier for inner dtypes in to_init_repr (#16235)assert_series_equal when categorical_as_str=True (#16700)read_database check for SQLAlchemy async Session objects (#16680)select and with_columns to idiomatic form (#16801)DataFrame.limit (#16753)include_nulls in DataFrame.update docstring (#16701)DataFrame.rolling (#16600)Expr/Series.map_elements (#16079)polars.sql docs entry and small docstring update (#16656)orient="row" in DataFrame constructor when applicable (#16977)Arc from FileCacheEntry (#16870)infer_schema_length as keyword-only argument in str.json_decode (#16835)ChunkedArray::from_chunks_and_dtype (#16697)1.0.0 release (#16705)Thank you to all our contributors for making this release possible! @JulianCologne, @KDruzhkin, @MarcoGorelli, @Object905, @alexander-beedie, @bertiewooster, @borchero, @coastalwhite, @datenzauberai, @dependabot, @dependabot[bot], @henryharbeck, @itamarst, @machow, @marenwestermann, @mcrumiller, @montanarograziano, @nameexhaustion, @orlp, @ritchie46, @siddharth-gulia, @stinodego, @tkellogg, @universalmind303 and @wence-
Remove deprecated parameters in Series.cut/qcut and update struct field names
reshape to return Array types instead of List types (#16825)get/gather operations (#16841)selector XOR set operation, guarantee consistent selector column-order (#16833)infer_schema_length as keyword-only argument in str.json_decode (#16835)set_sorted to only accept a single column (#16800)group_by iteration and partition_by to always return tuple keys (#16793)coalesce=False in left outer join (#16769)pyxlsb engine from read_database (#16784)Series.cut/qcut and update struct field names (#16741)top_k parameters nulls_last, maintain_order, and multithreaded (#16599)offset arg in truncate and round (#16655)offset in group_by_dynamic from 'negative every' to 'zero' (#16658)DataFrame.sql in favor of top-level pl.sql (#16598)Array instead of List (#16710)clip to no longer propagate nulls in the given bounds (#14413)str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)pivot when pivoting by multiple values (#16439)ewm_mean, ewm_std, and ewm_var (#15503)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)LazyFrame.with_context (#16860)descending to reverse in top_k methods (#16817)str.concat to str.join (#16790)arctan2d (#16786)split_at in split (#16865)split_at instead of double slice in chunk splits. (#16856)align_ if arrays are aligned (#16850)arg_sort (#16808)dt.offset_by 2x for constant durations (#16728)dt.truncate 1.5x faster when every is just a single duration (and not an expression) (#16666)read_csv SQL table reading function defaults (better handle dates) (#16866)VALUES clause and inline renaming of columns in CTE & derived table definitions (#16851)Enum values in lit (#16858).str.to_datetime when values are offset-aware (#16742)reshape to return Array types instead of List types (#16825)get/gather operations (#16841)SQL "SELECT" with no tables, optimise registration of globals (#16836)selector XOR set operation, guarantee consistent selector column-order (#16833)EXTRACT and DATE_PART SQL part abbreviations (#16767)set_sorted (#16800)coalesce=False in left outer join (#16769)read_database_uri passthrough from read_database (#16783)pyxlsb engine from read_database (#16784)check_order parameter to assert_series_equal (#16778)scan_csv (#16674)INTERVAL handling and improve related error messages, update sqlparser-rs lib (#16744)ORDER BY clause (#16745)pandas and pyarrow objects (#16746)Series.cut/qcut (#16741)date_range to no longer produce datetime ranges (#16734)min_periods as keyword-only for rolling methods (#16738)top_k parameters (#16599)NULLS FIRST/LAST ordering (#16711)INTERVAL strings (#16732)offset arg in truncate and round (#16655)offset in group_by_dynamic from "negative every" to "zero" (#16658)df.sql in favour of top-level pl.sql (#16598)clip bounds (#14413).str.to_datetime to default to microsecond precision for format specifiers "%f" and "%.f" (#13597)ewm_mean, ewm_std, and ewm_var (#15503)str.to_datetime (#16634)pl.read_json and DataFrame.write_json (#16550)nth to allow positional input of indices, remove columns parameter (#16510)rle output to len/value and update data type of len field (#15249)check_names parameter to Series.equals and default to False (#16610)SQLInterface and SQLSyntax errors (#16635)DIV function support to the SQL interface (#16678)should_rechunk check (#16852)read_excel and read_ods return identical frames across all engines when given empty spreadsheet tables (#16802)read_excel (#16840)top_k/bottom_k and fix a variety of bugs (#16804)DATE_PART SQL syntax/parsing, improve some error messages (#16761)pl. qualifier for inner dtypes in to_init_repr (#16235)assert_series_equal when categorical_as_str=True (#16700)read_database check for SQLAlchemy async Session objects (#16680)select and with_columns to idiomatic form (#16801)DataFrame.limit (#16753)include_nulls in DataFrame.update docstring (#16701)DataFrame.rolling (#16600)Expr/Series.map_elements (#16079)polars.sql docs entry and small docstring update (#16656)Arc from FileCacheEntry (#16870)infer_schema_length as keyword-only for str.json_decode (#16835)ChunkedArray::from_chunks_and_dtype (#16697)1.0.0 release (#16705)Thank you to all our contributors for making this release possible! @JulianCologne, @KDruzhkin, @MarcoGorelli, @Object905, @alexander-beedie, @bertiewooster, @coastalwhite, @datenzauberai, @dependabot, @dependabot[bot], @henryharbeck, @marenwestermann, @mcrumiller, @montanarograziano, @nameexhaustion, @orlp, @ritchie46, @siddharth-gulia, @stinodego, @universalmind303 and @wence-
> You can ignore the associated deprecation warning.
[!IMPORTANT]
The decision to change the default coalesce behavior of left join has been reversed. You can ignore the associated deprecation warning.
dtypes parameter to schema_overrides for read_csv/scan_csv/read_csv_batched (#16628)nulls_last/maintain_order/multithreaded parameters for top_k methods (#16597)SQLContext "eager_execution" param to "eager" (#16595)Series.equals parameter strict to check_dtypes and rename assertion utils parameter check_dtype to check_dtypes (#16573)DataFrame.serialize/deserialize (#16545)str.explode in favor of str.split("").explode() (#16508)nulls_last on sort operations (#16639)ARRAY literals and the UNNEST table function (#16330)struct.with_fields in grouping (#16629)TRY_CAST function (#16589)pl.sql function (#16528)DataFrame.serialize/deserialize (#16545)group_by_dynamic, upsample, and rolling (#16494)ORDER BY should not cause reordering of SELECT cols (#16579)shape in Array constructor and deprecate width parameter (#16567)Series in LazyFrame.select() (#16592)JOIN issues (#16507)sum over a list of strs (#16521)DataFrame.__getitem__ for empty list input - df[[]] (#16520)DataFrame.__getitem__ with 2 column inputs (#16517)LazyFrame properties may be expensive (#16618)versionadded tags, and add is_column_selection to the Expr meta docs (#16590)DataFrame.join docstring (#16576)implode reference from the user guide section on window functions (#16544)cargo update (#16574)typing.no_type_check (#16497)Thank you to all our contributors for making this release possible! @MarcoGorelli, @alexander-beedie, @coastalwhite, @hattajr, @itamarst, @mcrumiller, @nameexhaustion, @r-brink, @ritchie46, @stinodego, @twoertwein and @wence-
Add Series/Expr.has_nulls and deprecate Series.has_validity
Series/Expr.has_nulls and deprecate Series.has_validity (#16488)tree_format parameter for LazyFrame.explain in favor of format (#16486)DataFrame.__getitem__ improvements (#16495)is_column_selection() to expression meta, enhance expand_selector (#16479)Series/Expr.has_nulls and deprecate Series.has_validity (#16488)split_chunks for nested dtypes (#16493)top_k/bottom_k (#16489)COUNT(*) in SQL GROUP BY operations (#16465)nan_to_null when using multi-thread in pl.from_pandas (#16459)pl.field inside with_fields examples. (#16451)cum_max (#16456)Series/DataFrame.__getitem__ logic (#16482)Thank you to all our contributors for making this release possible! @BGR360, @alexander-beedie, @cmdlineluser, @coastalwhite, @itamarst, @marenwestermann, @mdavis-xyz, @messense, @orlp, @ritchie46 and @stinodego
Deprecate how="outer" join type in favour of how="full" (left/right are \*also\* outer joins)
how="outer" join type in favour of how="full" (left/right are *also* outer joins) (#16417)DataFrame.to_numpy (#16429)value_counts "count" column (#16434)alpha and alphanumeric selectors, add "ascii_only" to digit (#16362)__array__ method for Series and DataFrame to support copy parameter (#16401)read_excel dtype inference of "calamine" int/float results that include NaN (#16400)apply call in str_duration_ util. (#16412)interpolate_by entry to rst files. (#16422)Thank you to all our contributors for making this release possible! @KDruzhkin, @alexander-beedie, @ankane, @cmdlineluser, @coastalwhite, @itamarst, @nameexhaustion, @ritchie46 and @stinodego
Deprecate use_pyarrow parameter for to_numpy methods
use_pyarrow parameter for to_numpy methods (#16391)field expression as selector with an struct scope (#16402)DataFrame.to_numpy also for non-numeric frames (#16390)Series.to_numpy (#16383)DataFrame.to_numpy for Array/Struct types (#16386)DataFrame.to_numpy for Struct columns when structured=True (#16358)ClosedInterval in expr IR (#16369)to_numpy methods (#16394)Thank you to all our contributors for making this release possible! @MarcoGorelli, @alexander-beedie, @coastalwhite, @dangotbanned, @itamarst, @ritchie46, @stinodego and @wence-
use is\_sorted in ewm\_mean\_by, deprecate check\_sorted
[!WARNING]
This release was yanked. Please use the 0.20.28 release instead.
chunked to allow_chunks in parametric testing strategies (#16264)is_sorted for numeric data (#16333)Series.to_numpy performance for chunked Series that would otherwise be zero-copy (#16301)polars import (#16308)ctypes.util in CPU check script if possible (#16307)read_excel to handle bytes/BytesIO directly when using the "calamine" (fastexcel) engine (#16344)by column in rolling_*_by operations (#16249)Series.to_numpy (#16315)to_jax methods to support Jax Array export from DataFrame and Series (#16294)alpha, alphanumeric and digit selectors (#16310)require_all parameter to the by_name column selector (#15028)BytecodeParser for Python 3.13 (#16304)struct.with_fields (#16305)BETWEEN clause (#16279)cs.by_index, allow multiple indices for nth (#16217)excluded_dtypes list would grow indefinitely (#16340)map_elements typing (#16257)Series.reshape against invalid parameters (#16281)Series.to_numpy for Array types with nulls and nested Arrays (#16230)expand_selectors function, minor fixes (#16250)Object() (#16260)read_database overload (#16229)join docstring (#16299)DataFrame.to_numpy implementation to Rust side (#16354)interop::numpy module (#16346)DataFrame.to_numpy code (#16325)InterchangeDataFrame.version should be a ClassVar (not a property) (#16312)polars-expr README (#16316)cls (not self) in classmethods (#16303)Thank you to all our contributors for making this release possible! @MarcoGorelli, @NickCondron, @ShivMunagala, @alexander-beedie, @brandon-b-miller, @coastalwhite, @datenzauberai, @itamarst, @jsarbach, @max-muoto, @nameexhaustion, @orlp, @r-brink, @ritchie46, @stinodego, @thalassemia, @twoertwein and @wence-
Deprecate allow_infinities and null_probability args to parametric test strategies
allow_infinities and null_probability args to parametric test strategies (#16183)concat (#16128)to_torch "features" and "label" parameter behaviour when return type is not "dataset" (#16218)Enum types in parametric testing (#16188)GROUP BY ALL syntax and fix several issues with aliased group keys (#16179)write_database (#16099)nth(n) method, to go with existing first and last (#16112)dtype and strict in pl.Series's constructor for pyarrow arrays, numpy arrays, and pyarrow-backed pandas (#15962)IN clauses (#16101)Series functions (#16172)cumfold and cumreduce (#16173)to_numpy (#14353)Thank you to all our contributors for making this release possible! @MarcoGorelli, @YichiZhang0613, @alexander-beedie, @bertiewooster, @coastalwhite, @dangotbanned, @itamarst, @janpipek, @jrycw, @luke396, @nameexhaustion, @pydanny, @ritchie46, @stinodego, @thalassemia and @tharunsuresh-code
Revert "Add RLE to RLE_DICTIONARY encoder"
RLE_DICTIONARY encoder"ParameterCollisionError in read_excel (#16100)Thank you to all our contributors for making this release possible! @nameexhaustion, @ritchie46 and @wsyxbcl
> This release was yanked. Please use the 0.20.25 release instead.
[!WARNING]
This release was yanked. Please use the 0.20.25 release instead.
pytorch Tensor and Dataset export with new to_torch DataFrame/Series method (#15931)dd.mm.YYYY (#16045)pytorch Tensor and Dataset export with new to_torch DataFrame/Series method (#15931)uint datatype support for the SQL interface (#15993)NodeTraverser to Python (#15776)by argument for Expr.top_k and Expr.bottom_k (#15468)typed_lit to help schema determination in SQL "extract" func (#15955)pandas_to_pyseries function (#15948)read_csv_batched (#15944)fill_nan methods (pointing out that nan isn't null) (#16061)apply (#15982)sccache action (#16088)is_polars_dtype util (#16065)Thank you to all our contributors for making this release possible! @CanglongCl, @JulianCologne, @KDruzhkin, @MarcoGorelli, @alexander-beedie, @avimallu, @bertiewooster, @c-peters, @dependabot, @dependabot[bot], @eitsupi, @haocheng6, @itamarst, @luke396, @marenwestermann, @nameexhaustion, @orlp, @ritchie46, @stinodego, @thalassemia, @wence- and @wsyxbcl
Add a low-friction sql method for DataFrame and LazyFrame
sql method for DataFrame and LazyFrame (#15783)to_datetime (#15826)slope in interpolate (#15819)dt.round (#15861)is_not_nan (#15889)read_excel when using "calamine" engine (#15827)storage_options dict (take a shallow-copy) (#15859)shrink_dtype as non-streaming (#15828)ruff version and improve make clean on the Python side (#15858)rust-toolchain.toml from wheels (#15840)import_optional utility function (#15906)Thank you to all our contributors for making this release possible! @JulianCologne, @MarcoGorelli, @NedJWestern, @NexVeridian, @alexander-beedie, @deanm0000, @dependabot, @dependabot[bot], @ion-elgreco, @itamarst, @jr200, @nameexhaustion, @orlp, @reswqa, @ritchie46 and @stinodego
Various deprecation docstring improvements
read_excel and read_ods, use calamine engine for read_ods (#15808)read_database (#15809)read_excel and read_ods, use calamine engine for read_ods (#15808)dt.truncate supports broadcasting lhs (#15768)str.json_path_match (#15764)storage_options is passed to read_csv but fsspec isnt available (#15778)LazyFrame conversion errors (#15761)lit (#15718)read_parquet (#15770)is_between pushdown to scan_pyarrow_dataset (#15769)ewm_mean_by (#15687)prepare_expression_for_context shouldn't panic if exceptions raised from optimizer (#15681)Config.set_tbl_width_chars (#15566)Series/Expr.dt.truncate/round (#15698)json_path_match expr non-anonymous (#15682)Thank you to all our contributors for making this release possible! @MarcoGorelli, @NedJWestern, @Robinsane, @TobiasDummschat, @alexander-beedie, @c-peters, @dependabot, @dependabot[bot], @gasmith, @henryharbeck, @itamarst, @kszlim, @mbuhidar, @nameexhaustion, @orlp, @reswqa, @ritchie46, @stinodego and @wsyxbcl
Various deprecation docstring improvements
ewm_mean_by (#15687)prepare_expression_for_context shouldn't panic if exceptions raised from optimizer (#15681)json_path_match expr non-anonymous (#15682)Thank you to all our contributors for making this release possible! @henryharbeck, @reswqa and @ritchie46
Add missing deprecation warning to DataFrame.replace
group_by multiple null columns produce phantom row (#15659)arr.min/max (#15654)list.mean fast path shouldn't produce NaN (#15652)DataFrame.replace (#15612)offset deprecation in upsample (#15636)Thank you to all our contributors for making this release possible! @MarcoGorelli, @Priyansh4444, @StevenMia, @eitsupi, @itamarst, @mcrumiller, @orlp, @reswqa, @ritchie46 and @stinodego
Replace most deprecated calls with bounded version
Filter,Select,WithColumns (#15608)AnyValue (#15576)str.head and str.tail (#14425)union/or operator for pl.Enum (#14965)BytecodeParser to handle additional math functions, and imports from the global namespace (#15627)is_between expressions to Arrow (#15180)to_integer (#15604)null_on_oob parameter to expr.array.get (#15426)is_first/last_distinct for not nested non-numeric list (#15552)mean and median (#14471)write_excel that could lead to incorrect spanning range determination (#15631)mean_horizontal on a single column (#15118)AggregatedScalar (#15606)GROUP BY clauses that use position ordinals (#15584)sort with SortOptions and SortMultipleOptions (#15590)Thank you to all our contributors for making this release possible! @CanglongCl, @ChayimFriedman2, @Fokko, @JamesCE2001, @MarcoGorelli, @NedJWestern, @TrevorWinstral, @alexander-beedie, @deanm0000, @douglas-raillard-arm, @eitsupi, @filabrazilska, @i-aki-y, @itamarst, @leoforney, @mcrumiller, @nameexhaustion, @orlp, @ozgrakkurt, @reswqa, @ritchie46 and @stinodego
Replace std::thread spawn with tokio block\_in\_place
MEDIAN aggfunc (#15519)string, boolean and binary dtype in top_k (#15488)TRUNCATE TABLE command (#15513)GREATEST and LEAST (#15511)read/scan_parquet (#15434)agg_list for NullChunked (#15439)skip_rows_after_header to pyarrow csv reader (#15533)schema_overrides contains nonexistent columns (#15528)list.get should take validity into account (#15516)group_by partitioned with literal Series panic (#15487)GroupsProxy::Slice windows (#15509)pow return type evaluation (#15506)read_database draining iter_batches early (#15504).filter() (#15445)n into clear (#15432)by parameter to group_by in DataFrame/LazyFrame.upsample/group_by_dynamic/rolling (#15527)make docs command, DataType docs/layout tweak, minor README updates (#15386)Series.list.median. (#15451)read_parquet (#15532)DataFrame._read classmethods (#15521)io.database executor module (#15526)hive_schema functionality (#15508)Thank you to all our contributors for making this release possible! @CanglongCl, @ChayimFriedman2, @MarcoGorelli, @alexander-beedie, @cmdlineluser, @dependabot, @dependabot[bot], @henryharbeck, @mbuhidar, @nameexhaustion, @reswqa, @ritchie46, @rob-sil and @stinodego
CSV reading memory usage tests and fixes
explode_by_offsets for decimal (#15417)read_clipboard and DataFrame.write_clipboard (#15272)null_on_oob parameter to expr.list.get (#15395)n_unique() in group-by context when group is empty (#15289)to_any_value should supports all LiteralValue type (#15387)sort for series with unsupported dtype should raise instead of panic (#15385)explode mapping strategy in pl.Expr.over (#15402)outer_coalesce join strategy in the user guide (#15405)series/array.py (#15383)arg_sort and arg_sort_by (#15348)Thank you to all our contributors for making this release possible! @CanglongCl, @JamesCE2001, @MarcoGorelli, @Sol-Hee, @alexander-beedie, @dependabot, @dependabot[bot], @itamarst, @kszlim, @mcrumiller, @nameexhaustion, @orlp, @reswqa, @ritchie46, @rob-sil and @thomaslin2020
Rename parameter by to group_by in DataFrame.upsample/group_by_dynamic/rolling
by to group_by in DataFrame.upsample/group_by_dynamic/rolling (#14840)from_repr parameter from tbl to data (#15156)arr.n_unique (#15296)read_database support for SurrealDB ("ws" and "http") (#15269)Sequence in from_records (#15329)async database calls (#15202)name parameter to GroupBy.len method (#15235)read_database when reading from Kùzu graph database (#15218)map_elements is called without return_dtype specified (#15188)async SQLAlchemy connections to read_database (#15162)time_unit in pl.duration when nanoseconds is specified (#14987)strict parameter to from_dict/from_records (#15158)s.clear() when dtype is Object (#15315)Series.list.std and Series.list.var (#15267)LazyFrame (#15331)from_dicts (#15344)schema_overrides contains nonexistent column name (#15290)dtype input for int_range and int_ranges (#15339)LazyFrame (#15297)strict flag when constructing a Struct Series from any values (#15302)DataFrame init from dict (#15217)check_sorted in some cases (#15227)rle expression (#15248)read_parquet when columns parameter is specified (#15229)cs.temporal() selector uses wildcard time zone matching for Datetime (#13683)TypeError on constructor failure (#15178)timestamp example (#15281)Series.search_sorted (#14737)is_between, and add an example (#15197)clear operation (#15304)Cache[count] to Cache[cache_hits] (#15300)PyDataFrame.from_dicts (#15274)wrapping_abs to arithmetic kernel (#15210)RUST_BACKTRACE=1 in the CI test suite (#15204)read_database functionality into cleaner module structure (#15201)dataframe module in PyO3 bindings (#15165)Thank you to all our contributors for making this release possible! @MarcoGorelli, @alexander-beedie, @braaannigan, @c-peters, @cojmeister, @deanm0000, @dependabot, @dependabot[bot], @itamarst, @kszlim, @mbuhidar, @mcrumiller, @mickvangelderen, @orlp, @petrosbar, @reswqa, @ritchie46, @rob-sil, @sportfloh, @stinodego and @yutannihilation
add new when-then-otherwise kernels
(first|last)_non_null (#15050)read_database results (#15126)closed and by are passed to rolling_* aggregations (#15108)by of invalid dtype (#15088)non_existent arg to replace_time_zone (#15062)register_plugin a standalone function and include shared lib discovery (#14804)infer_schema_length parameter on read_database (#15076)strict parameter to DataFrame constructor to allow non-strict construction (#15034)u32 when sum_horizontal provided with single boolean column (#15114)product on an invalid type (#15093)max() on sorted float arrays if it exists instead of NaN (#15060)nulls_last in streaming sort (#15061)count agg (#15051)string_addition_to_linear_concat (#15006)new and old parameters in replace description (#15019)Thank you to all our contributors for making this release possible! @JackRolfe, @MKisilyov, @MarcoGorelli, @alexander-beedie, @c-peters, @flisky, @jqnatividad, @mcrumiller, @mickvangelderen, @nameexhaustion, @orlp, @petrosbar, @ritchie46, @stinodego and @trueb2
Use sorted flag for (first|last)_non_null
(first|last)_non_null (#15050)strict parameter to DataFrame constructor to allow non-strict construction (#15034)nulls_last in streaming sort (#15061)count agg (#15051)string_addition_to_linear_concat (#15006)new and old parameters in replace description (#15019)Thank you to all our contributors for making this release possible! @MKisilyov, @MarcoGorelli, @alexander-beedie, @c-peters, @flisky, @jqnatividad, @mcrumiller, @mickvangelderen, @nameexhaustion, @petrosbar, @ritchie46, @stinodego and @trueb2
Ensure parallel encoding/compression in sink_parquet
sink_parquet (#14964)Array type in parquet (#14943)drop_first parameter to Series.to_dummies (#14846)read_database_uri (#14682)count_rows multi-threaded under-counting in parser.rs (#14963)read_database behaviour with empty ODBC "iter_batches" (#14918)RecordBatch objects when pyarrow <= 12 (#14922)read_database now properly handles empty result sets from arrow-odbc (#14916)with_columns (#14859)include_index in from_pandas regarding "default indices" (#14920)fastexcel (#14907)POLARS_FORCE_ASYNC env var parsing (#14909)Thank you to all our contributors for making this release possible! @MarcoGorelli, @alexander-beedie, @ambidextrous, @battmdpkq, @mcrumiller, @mickvangelderen, @orlp, @petrosbar and @ritchie46
Deprecate overwrite_schema parameter for DataFrame.write_delta
overwrite_schema parameter for DataFrame.write_delta (#14879)cum_count on columns (#14849)__slots__ to Polars classes (#14857)fastexcel to show_versions (#14869)cum_count on columns (#14849)pl.read_database (#14822)DataFrame.min/max for decimals (#14890)__reduce__ implementation on DataType object (#14778)ambiguous instead of use_earliest (#14820)asof from join strategy, change parameter from strategy to how in user guide (#14793)utils module to _utils to explicitly mark it as private (#14772)_cpu_check module (#14768)Thank you to all our contributors for making this release possible! @MarcoGorelli, @Sol-Hee, @alexander-beedie, @c-peters, @deanm0000, @dependabot, @dependabot[bot], @eitsupi, @flisky, @geekvest, @mcrumiller, @mickvangelderen, @nameexhaustion, @orlp, @petrosbar, @ritchie46 and @stinodego
Elide utf8/binary cast in Parquet reading
_cpu_check ("read_cpu_flags") (#14758)Thank you to all our contributors for making this release possible! @alexander-beedie, @ritchie46 and @stinodego
Add allow_copy parameter to DataFrame.to_numpy
allow_copy parameter to DataFrame.to_numpy (#14569)from_arrow when array has 0 chunks (#14562)map_elements for additional string functions (#14565)allow_copy parameter to DataFrame.to_numpy (#14569)read_database interop with sqlalchemy Session and various Result objects (#14557)map_elements for temporal attributes/methods (#14529)Thank you to all our contributors for making this release possible! @CBell045, @alexander-beedie, @c-peters, @nameexhaustion, @ritchie46 and @stinodego
use owned arithmetic in horizontal\_sum
writable flag to DataFrame.to_numpy (#14520)flush operator to streaming operators (#14500)transpose (#14527)is_numeric check on Series.std/var (#14493)schema input in DataFrame constructor (#14483)Config.save_to_file (#14533)infer_schema_length param description (#14233)README.md (#14488)Series.bin.ends_with, Series.bin.starts_with, Series.bin.decode, Series.bin.encode. (#14478)Thank you to all our contributors for making this release possible! @FBruzzesi, @NedJWestern, @c-peters, @dannyfriar, @i-aki-y, @jdanford, @mbuhidar, @mcrumiller, @ritchie46, @stinodego and @taki-mekhalfa
Deprecate positional args in pivot to prepare new functionality
pivot to prepare new functionality (#14428)pivot would introduce duplicate column names (#14431)min/max for categorical dtype (#14112)polars.testing.* in pytest stack traces (#14399)clip (#14410)mean_horizontal expression (#14369)arr.shift (#14298)list.n_unique (#14306)Categorical/Enum in Series.to_numpy (#14275)Array dtype (#14265)mean and median (#14376)index and one of them was Struct (#14438)Series from projection state (#14437)index was Struct (#14308)clip inputs (#14416)map_elements (#14397)Series.to_numpy (#14341)set_operation if the input is sliced and be broadcast (#14303)par_iter in list.to_struct by POOL.install (#14304)list.get does not work on list of decimals (#14276)read_database docstring note about getting the connection URI string for sqlalchemy (#14461)slice / len_chars (#14395)DataFrame.rows_by_key (#14149)Series.bin.contains (#14297)arg_min/max test case (#14439)setup-graphviz action to v2 (#14418)make clean command (#14408)_or to or_ in PyO3 (same for _xor/_and) (#14393)DataFrame.to_numpy structured code (#14348)Series.to_numpy to handle Decimal/Time types in Rust (#14296)Series.to_numpy with timezones (#14337)Thank you to all our contributors for making this release possible! @BGR360, @CaselIT, @MarcoGorelli, @Migi, @NedJWestern, @Vincenthays, @alexander-beedie, @deanm0000, @dependabot, @dependabot[bot], @engdoreis, @flisky, @grinya007, @itamarst, @janosh, @kalekundert, @lukemanley, @mbuhidar, @mcrumiller, @petrosbar, @r-brink, @rben01, @reswqa, @ritchie46, @stinodego, @taki-mekhalfa and @thomasfrederikhoeck
Improve deprecation message of dtype_if_empty param
threadpool_size to thread_pool_size (#14236)is_not_null is used (#14260)Series.to_numpy for boolean/temporal types (#14261)UnitVec in polars-plan traversal (#14199)UnitVec in streaming joins (#14197)ChunkId (#14175)u8/i8/u16/i16 parsers to CSV reader (#14241)F-order data in and out of numpy to polars zero copy (#14259)list.gather_every (#14253)prefix/suffix_fields (#14251)Series.to_numpy to return f64 for Int32/UInt32 Series with nulls instead of f32 (#14240)read_excel format detection, and support for excel 97-2004 workbooks (#14234)arr.to_struct (#14202)IdxVec generic as UnitVec (#14196)unique and hash_rows for null column (#14111)Null columns (#14107)group_by (#14071)list & array measures of dispersion (#13245)convert_time_zone on time-zone-naive datetime, convert as if converting from UTC (#13960)glimpse overload signature (#14258)any/all_horizontal with single input has incorrect type (#14256)Series.to_numpy on booleans without nulls return bool type (#14239)is_elementwise=True parameter) (#14135)read_excel "calamine" (fastexcel) engine (#14171)set_operations of binary dtype (#14152)Date to Time and vice versa (#14127)gt/lt cmp for null dtype (#14119)read_excel updates (#14039)1970-01-01 (#14050)Object dtype designation (#14072)melt panic when there are no value vars (#14057)json_encode should respect the logical type (#14063)SliceSink with empty data (#14025)Series.to_pandas for categorical types (#14028)any_horizontal and all_horizontal (#14148)return_dtype parameter for map_elements and map_batches (#14114)str.replace and str.replace_all (#13382)name operation is allowed per expression (#14075)dtype_if_empty param (#14068)*args in Series.to_numpy (#14248)meta module (#14230).cargo directory to .gitignore (#14191)take_chunked to polars-ops (#14185)cargo update (#14160)Thank you to all our contributors for making this release possible! @JulianCologne, @MarcoGorelli, @Vincenthays, @Wainberg, @alexander-beedie, @apcamargo, @braaannigan, @c-peters, @deanm0000, @dependabot, @dependabot[bot], @dpinol, @edavisau, @eitsupi, @flisky, @grinya007, @ion-elgreco, @itamarst, @lukemanley, @mcrumiller, @orlp, @r-brink, @reswqa, @ritchie46, @stinodego and @taki-mekhalfa
Deprecate dtype_if_empty parameter for Series constructor
String/Binary type. (#13748)dtype_if_empty parameter for Series constructor (#13976)read_excel, using fastexcel (~8-10x speedup) (#14000)DataFrame.describe by presorting columns (#13822)arr.sum for inner non-null bool (#13800)UnstableWarning for unstable functionality (#13948)read_excel, using fastexcel (~8-10x speedup) (#14000)describe on a LazyFrame (#13982)timestamp precision modifier (#13936)LEFT, RIGHT and SUBSTR SQL string funcs (#13888)explode for ArrayNameSpace (#13923)describe code (#13720)ignore_nulls for arr.join (#13919)ignore_nulls for list.join (#13701)ignore_nulls for pl.concat_str (#13877)int_range and int_ranges signatures (#13867)binview (#13871)-pl.col(...) (#13776)EXTRACT with "century", "millennium", and "timezone" parts (#13634)numeric and/or decimal types (#13739)str.zfill (#13790)String/Binary type. (#13748)nulls_last for Series.sort (#13794)ftp URLs, improve URL check (#13781)list.min/max with empty and/or None elements (#14018)to_pandas() work for Dataframe and Series with dtype Object (#13910)pl.concat(how="align") when no columns are shared between frames (#13941)date_as_object=False as default for Series.to_pandas (just like DataFrame.to_pandas) (#13984)arr/list.contains (#13959)max_colname_length formatting in glimpse() (#13969)is_not_null for Struct columns (#13921)sum_horizontal (#13880)min_periods (#13863)ignore_nulls for list.join (#13701)filter behaviour matches docstring (expect equivalence with eq) (#13864)abs for Decimal, error on Date/Time/Datetime (#13821)gather_every should work on agg context (#13810)is_in (#13814)None, fix Series edge-case (#13780)glimpse test (#13979)describe tidy-up, and slight rewording of some Exception docstrings (#13942)filter to polars-compute (#13897)expr/general non-anonymous (#13832)Thank you to all our contributors for making this release possible! @ByteNybbler, @JulianCologne, @MarcoGorelli, @Wainberg, @alexander-beedie, @c-peters, @dependabot, @dependabot[bot], @edavisau, @flisky, @ion-elgreco, @itamarst, @jacksonthall22, @kstoneriv3, @mcrumiller, @mkucijan, @nameexhaustion, @orlp, @petrosbar, @r-brink, @reswqa, @ritchie46, @stinodego, @taki-mekhalfa, @thomasaarholt and @valorien
new implementation for String/Binary type.
String/Binary type. (#13748)arr.sum for inner non-null bool (#13800)ignore_nulls for list.join (#13701)ignore_nulls for pl.concat_str (#13877)int_range and int_ranges signatures (#13867)binview (#13871)-pl.col(...) (#13776)EXTRACT with "century", "millennium", and "timezone" parts (#13634)numeric and/or decimal types (#13739)str.zfill (#13790)String/Binary type. (#13748)nulls_last for Series.sort (#13794)ftp URLs, improve URL check (#13781)sum_horizontal (#13880)min_periods (#13863)ignore_nulls for list.join (#13701)filter behaviour matches docstring (expect equivalence with eq) (#13864)abs for Decimal, error on Date/Time/Datetime (#13821)gather_every should work on agg context (#13810)is_in (#13814)None, fix Series edge-case (#13780)filter to polars-compute (#13897)expr/general non-anonymous (#13832)Thank you to all our contributors for making this release possible! @ByteNybbler, @MarcoGorelli, @Wainberg, @alexander-beedie, @dependabot, @dependabot[bot], @edavisau, @flisky, @ion-elgreco, @itamarst, @kstoneriv3, @mcrumiller, @mkucijan, @nameexhaustion, @orlp, @reswqa, @ritchie46, @stinodego, @taki-mekhalfa and @thomasaarholt
Your coding agent can read these notes before it upgrades. Set up the MCP server →