NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #357 most downloaded on PyPI
Pandas on AWS.
Last release 1 months ago
03 Aug 2026
Ships fairly regularly
a new release about every 2 months
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
8 years old
167 releases · first in 2019
> ⚠️ For platforms without PyArrow 3 support (e.g. MWAA, [EMR](https://aws-data-wrangler.readthedocs.io/en/stable/install.html#emr-cluster), [Glue PyS
⚠️ For platforms without PyArrow 3 support (e.g. MWAA, EMR, Glue PySpark Job):<br> ➡️
pip install pyarrow==2 awswrangler
ExpectedBucketOwner #562We thank the following contributors/users for their work on this release:
@maxispeicher, @impredicative, @adarsh-chauhan, @Malkard.
P.S. The AWS Lambda Layer file (.zip) and the AWS Glue file (.whl) are available below. Just upload it and run!
Redshift COPY now supports the new SUPER type (i.e. SERIALIZETOJSON) #514
One column per quarter.
We thank the following contributors/users for their work on this release:
@maxispeicher, @danielwo, @jiteshsoni, @igorborgest, @njdanielsen, @eric-valente, @gvermillion, @zseder, @gdbassett, @orenmazor, @senorkrabs, @Natalie-Caruana.
P.S. The AWS Lambda Layer file (.zip) and the AWS Glue file (.whl) are available below. Just upload it and run!
SQLServer support (Driver must be installed separately) #356
s3.read_parquet_table() #495keep_files behavior for failed Redshift COPY executions #505We thank the following contributors/users for their work on this release:
@maxispeicher, @danielwo, @jiteshsoni, @gvermillion, @rodalarcon, @imanebosch, @dwbelliston, @tochandrashekhar, @kylepierce, @njdanielsen, @jasadams, @gtossou, @JasonSanchez, @kokes, @hanan-vian @igorborgest.
P.S. The AWS Lambda Layer file (.zip) and the AWS Glue file (.whl) are available below. Just upload it and run!
Add aws_access_key_id, aws_secret_access_key, aws_session_token and boto3_session for Redshift copy/unload #484
aws_access_key_id, aws_secret_access_key, aws_session_token and boto3_session for Redshift copy/unload #484We thank the following contributors/users for their work on this release:
@danielwo, @thetimbecker, @njdanielsen, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Add secretmanager module and support for databases connections #402
con = wr.redshift.connect(secret_id="my-secret", dbname="my-db")
df = wr.redshift.read_sql_query("SELECT ...", con=con)
con.close()
wr.*.connect() #481We thank the following contributors/users for their work on this release:
@danielwo, @nmduarteus, @nivf33, @kinghuang, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
New wr.timestream.create_database() function
ignore_empty argument to ignore 0 bytes files for:
We thank the following contributors/users for their work on this release:
@danielwo, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
sqlalchemy and psycopg2 dependencies replaced by `redshift_connector` and pg8000
sqlalchemy and psycopg2 dependencies replaced by redshift_connector and pg8000wr.db.* functions was distributed into wr.redshift.*, wr.postgresql.* and wr.mysql.* (Tutorial)wr.redshift.* (Tutorial)wr.catalog.get_engine() was replaced by wr.redshift.connect(), wr.postgresql.connect(), wr.mysql.connect() (Tutorial)We thank the following contributors/users for their work on this release:
@Brooke-white, @danielwo, @sapientderek, @pmleveque, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Deterministic result for s3.read_parquet_metadata() #449
catalog.add_column() #451catalog.delete_column() #451s3.read_parquet_metadata() #449ctas_approach=False and chunksize=True #458We thank the following contributors/users for their work on this release:
@tuannguyen0901, @bryanyang0528, @czagoni, @jesusch, @danielwo, @DonghanYang, @eric-valente, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Add configurable Endpoint URL for AWS services #418
wr.db.read_sql_query() #431wr.db.read_sql_query() #427wr.s3.to_csv() running on Windows platform.We thank the following contributors/users for their work on this release:
@martinSpears-ECS, @imanebosch, @Eric-He-98, @brombach, @Thomas-Hirsch, @vuchetichbalint, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Add encrypted glue connection management #413
s3.to_csv()) #415s3.read_parquet failing with some timezone aware columns #417We thank the following contributors/users for their work on this release:
@jeanbaptistepriez, @mike-at-upside, @Thiago-Dantas, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
General exceptions handling improvements #409
We thank the following contributors/users for their work on this release:
@tasq-inc, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Add s3_additional_kwargs for wr.s3.copy_objects() and wr.s3.merge_datasets() #388
s3_additional_kwargs for wr.s3.copy_objects() and wr.s3.merge_datasets() #388data_source argument for Athena queries #392tinyint columns on Redshift loads #400wr.s3.to_parquet() calls #399We thank the following contributors/users for their work on this release:
@timgates42, @bvsubhash, @DonghanYang, @sl-antoinelaborde, @Xiangyu-C, @tuannguyen0901, @JPFrancoia, @sapientderek, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Fix bug for wr.s3.read_parquet() with timezone offset. #385
wr.s3.read_parquet() with timezone offset. #385We thank the following contributors/users for their work on this release:
@chrisrana, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Fix issues in reading Parquet files with timestamp (timezone aware) columns. #382 #383
We thank the following contributors/users for their work on this release:
@tasq-inc, @chrisrana, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Significant Amazon S3 I/O speed up for big files #377
We thank the following contributors/users for their work on this release:
@jarretg, @chrisrana, @vikramshitole, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Global configuration s3fs_block_size was replaced by s3_block_size #370
s3fs_block_size was replaced by s3_block_size #370schema_evolution argument. #353s3fs dependency was replaced by builtin code. #370We thank the following contributors/users for their work on this release:
@isrsal, @bppont, @weishao-aws, @alexifm, @Digma, @samcon, @TerrellV, @msantino, @alvaropc, @luigift, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel/egg files are available below. Just upload it and run!
Fix NaN values handling for wr.athena.read_sql_*(). #351
wr.athena.read_sql_*(). #351We thank the following contributors/users for their work on this release:
@czagoni, @josecw, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel file are available below. Just upload it and run!
wr.s3.to_parquet() now has max_rows_by_file argument. #283
wr.s3.to_parquet() now has max_rows_by_file argument. #283*, ?, [seq], [!seq]) for any list/read/delete/copy function on S3. #322wr.s3.to_parquet() during appends. #342wr.s3.to_parquet/csv(). #343We thank the following contributors/users for their work on this release:
@Thiago-Dantas, @andre-marcos-perez, @ericct, @marcelo-vilela, @edvorkin, @nicholas-miles, @chrispruitt, @rparthas ,@igorborgest.
P.S. Lambda Layer zip file and Glue wheel file are available below. Just upload it and run!
The partitioned parquet reading now has a different approach for pushdown filters. For details check the tutorial
wr.athane.read_sql_*() - TUTORIAL #331wr.athena.describe_table() #329wr.athena.show_create_table() #334path_ignore_suffix argument to all read functions #326PyArrow 1.0.0 #337Pandas 1.1.0wr.athane.read_sql_*() now accepts empty results #299skip_header_line_count argument to wr.catalog.create_csv_table() #338wr.s3.read_csv() slow with chunksize #324wr.s3.read_csv() with "chunksize" does not forward pandas_kwargs "encoding" #330wr.athane.read_sql_*() w/ ctas_approach=True #335We thank the following contributors/users for their work on this release:
@kylepierce, @davidszotten, @meganburger, @erikcw, @JPFrancoia, @zacharycarter, @DavideBossoli88, @c-line, @anand086, @jasadams, @mrtns, @schot, @koiker, @flaviomax, @bryanyang0528, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel file are available below. Just upload it and run!
Add wr.catalog.get_partitions(). #305
wr.catalog.get_partitions(). #305We thank the following contributors/users for their work on this release:
@jasadams, @bryanyang0528, @qemtek, @igorborgest.
P.S. Lambda Layer zip file and Glue wheel file are available below. Just upload it and run!
Now casting columns before append on an existing table only if necessary (wr.s3.to_parquet()).
wr.s3.to_parquet()).flag.writeable==False)P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Casting support for *any* column type to string using dtype argument on wr.s3.to_parquet()
dtype argument on wr.s3.to_parquet()P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add auto_create and db_groups arguments to get_redshift_temp_engine #288
auto_create and db_groups arguments to get_redshift_temp_engine #288validate_schema arguments to wr.s3.read_parquet_tablesafe argument to read_parquet #296last_modified_begin and last_modified_begin to list_objects, read_csv, read_json, read_fwf and read_parquetget_table_description on tables w/o description #294We thank the following contributors/users for their work on this release:
@koiker, @patrick-muller, @flaviomax, @acere, @jarretg, @bryanyang0528, @schrobot, @kinghuang, @igorborgest.
P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add create/delete database on wr.glue
sanitize_columns arg for s3.to_parquet and s3.to_csv #278 #279We thank the following contributors/users for their work on this release:
@ywang103, @patrick-muller, @tuliocasagrande, @sarojdongol, @sdknij, @ilyanoskov, @igorborgest.
P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add support for reading CSV, JSON and FWF partitions. #265
encoding arg support for reading CSV, JSON and FWF. #271We thank the following contributors/users for their work on this release:
@bryanyang0528, @dwbelliston, @patrick-muller, @sdknij, @igorborgest.
P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Support for Athena Partition Projection [TUTORIAL]
1.3.15 #259dtype (cast) on wr.s3.to_parquet with nested types #263us-east-1 #252wr.s3.to_parquet for partitions in reverse order #264We thank the following contributors/users for their work on this release:
@bryanyang0528, @zachmoshe, @buseynehannes, @jiajie999, @igorborgest.
P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Infer mixed Parquet schemas on wr.s3.read_parquet_metadata and wr.s3.store_parquet_metadata #195
s3fs version bumped #236We thank the following contributors/users for their work on this release:
@mrshu, @bryanyang0528, @JPFrancoia, @jaidisido, @qemtek, @dwbelliston, @mbiemann, @parasml, @BrainMonkey, @hyperloglog, @igorborgest.
P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add support for uint8, uint16, uint32 and uint64 on Parquet. #76
uint8, uint16, uint32 and uint64 on Parquet. #76get_table_parameters, upsert_table_parameters and upsert_table_parameters on wr.catalog. #224cache for s3fs.s3.to_parquet overwriting with different partition schema.We thank the following contributors/users for their work on this release:
@robertaves ,@jar-no1, @JPFrancoia, @igorborgest.
P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Removing objects ending with "/" from wr.s3.list_objects()
wr.s3.list_objects()P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Support for nested arrays and structs on wr.s3.to_parquet() #206
wr.s3.to_parquet() #206custom_classifications to wr.emr.create_cluster() #193kms_key_id, max_file_size, region arguments to wr.db.unload_redshift() #197catalog_versioning argument to wr.s3.to_csv() and wr.s3.to_parquet() #198keep_files and ctas_temp_table_name arguments to wr.athena.read_sql_*() #203replace_filenames argument to wr.s3.copy_objects() #215wr.s3.to_csv() and wr.s3.to_parquet() no longer need delete table permission to overwrite catalog table #198wr.db.read_sql_query()(PostgreSQL) #200We thank the following contributors/users for their work on this release:
@robkano ,@luigift, @parasml, @OElesin, @jar-no1, @keatmin, @pmleveque, @sapientderek, @jadayn, @igorborgest.
P.S. Lambda Layer's zip-file and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add wr.s3.copy_objects and wr.s3.merge_datasets #186
append mode for wr.catalog.create_parquet_table and wr.catalog.create_csv_table #188We thank the following contributors/users for their work on this release:
@JPFrancoia, @deathrowe, @igorborgest.
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add moto support for S3 and EMR (*partially*) #109
We thank the following contributors/users for their work on this release:
@russellbrooks, @vincentclaes, @JPFrancoia, @igorborgest.
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add new wr.catalog.extract_athena_types() #141
validate_schema to wr.s3.to_parquet() #167boto3_session #172We thank the following contributors/users for their work on this release:
@vfrank66, @JPFrancoia, @jewelltp, @hjuhel-cdpq, @jar-no1, @rmlove, @josecw, @igorborgest.
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
categories arg in s3.read_parquet, db.unload_redshift, athena.read_sql_query [#160]
categories arg in s3.read_parquet, db.unload_redshift, athena.read_sql_query [#160]We thank the following contributors/users for their work on this release:
@vfrank66, @nitin-kakkar, @sapientderek, @nagomiso, @igorborgest.
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Check out the brand new documentation page!
Check out the brand new documentation page!
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add header and filename arguments to Pandas.to_csv()
header and filename arguments to Pandas.to_csv()P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add pandas.read_fwf(), read_fwf_list(), read_fwf_prefix() for fixed-width files #131
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Support for all pandas.read_csv() arguments
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Support for categorical partitions for Pandas.to_parquet() #115
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Pandas.to_aurora() improvements
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Nothing published for this version
Nothing published for this version
Nothing published for this version
Support for empty dataframe for Pandas.read_sql_athena(ctas_approach=True)
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Add description, parameters and column's comments as arguments to all methods that creates any Glue tables (METADATA).
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Pandas -> Aurora (MySQL/PostgreSQL) (Append/Overwrite) (Via S3)
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
P.P.S. Have you never used Layers? Check the step-by-step guide.
P.P.P.S. AWS Data Wrangler counts on compiled dependencies (C/C++) so there is no support for Glue PySpark by now (Only Glue Python Shell).
Fix Default Session bug for environments without credentials
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
Nothing published for this version
Pandas to Redshift with upsert mode
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
Read Parquet tables from Glue Catalog directly to Pandas DataFrame
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
Read parquet data from s3 directly to Pandas DataFrame #73
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. Just upload it and run!
Add support for Decimal data type #58
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Improving cast for date columns
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Setting null date values as None for pandas.read_sql_athena() #69
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Fix bug for boolean type on spark.to_redshift()
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
P.S.* Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!*
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Fix issues for partitions with a single raw #62
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Fix cast issues for partition columns #59 #60
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Pandas.read_sql_athena() now also accept bool columns with null values.
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Add *tags* parameter to EMR.create_cluster()
P.S. Lambda Layer's bundle and Glue's wheel/egg are available below. It's just upload and run!
Your coding agent can read these notes before it upgrades. Set up the MCP server →