NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1087 most downloaded on PyPI
Snowflake Snowpark for Python
Last release 24 days ago
10 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
4 versions withdrawn
withdrawn after publishing
4 years old
76 releases · first in 2022
One column per quarter.
Added DataFrame.to_polars() to convert a Snowpark DataFrame to a Polars DataFrame. Supports an optional transport="parquet" mode for large data-transf
DataFrame.to_polars() to convert a Snowpark DataFrame to a Polars DataFrame. Supports an optional transport="parquet" mode for large data-transfer-dominated workloads.datetime.timedelta as the type annotation for day-time interval (DayTimeIntervalType) parameters and return values, and YearMonthInterval (a type annotation sentinel from snowflake.snowpark.types) for year-month interval (YearMonthIntervalType) parameters and return values.table_properties key in the iceberg_config dictionary of DataFrameWriter.save_as_table, which emits a TABLE_PROPERTIES = ('k'='v', ...) clause on Iceberg table creation (CREATE / CTAS).SQL compilation error: invalid default argument expression. Integer defaults were emitted as DEFAULT <value> :: INT, and a parameter default only accepts a constant expression. The redundant cast is no longer generated.response_format shape for ai_extract and DataFrame.ai.extract.test_generator_table_function failing as assert 0 < 0 on the seq2(0) + generator(timelimit => 1) case. seq2(0) passes sign=0, so it continues at 0 after 32767, and a one-second generator run usually emits far more than one 32768-value cycle, making ORDER BY seq2(0) LIMIT 3 return 0, 0, 0. That assertion now allows equal values; the seq1(1) + rowcount assertions stay strictly increasing.modin.pandas inside a Snowpark pandas apply() function raised a ModuleNotFoundError.Fixed DataFrame.ai.count_tokens to call the AI_COUNT_TOKENS SQL function instead of the deprecated SNOWFLAKE.CORTEX.COUNT_TOKENS. Token counts may dif…
ai_count_tokens function to snowflake.snowpark.functions to estimate token counts for AI function calls.ai_multi_embed function to snowflake.snowpark.functions to generate multimodal embeddings from text, images, audio, or video files.ai_redact function to snowflake.snowpark.functions to detect and redact personally identifiable information (PII) from text.DataFrame.ai.multi_embed method to generate multimodal embeddings via the DataFrame API.DataFrame.ai.redact method to detect and redact PII from text columns via the DataFrame API.DataFrame.ai.translate method to translate text columns between languages via the DataFrame API.experimental tag from all AI SQL functions in DataFrameAIFunctions (complete, filter, agg, classify, similarity, sentiment, embed, summarize_agg, transcribe, parse_document, extract, count_tokens, split_text_markdown_header, split_text_recursive_character) and RelationalGroupedDataFrame.ai_agg.DataFrame.ai.count_tokens to match the standalone ai_count_tokens function: function_name is now the required first parameter, and model is optional. Also added options and return_error_details parameters.DataFrame.ai.count_tokens to call the AI_COUNT_TOKENS SQL function instead of the deprecated SNOWFLAKE.CORTEX.COUNT_TOKENS. Token counts may differ slightly due to the updated tokenizer.SELECT * from joins, which was causing performance regressions. This has no functional impact.No user-facing changes in this release.
No user-facing changes in this release.
Added the udf_init_once decorator in snowflake.snowpark.functions for marking functions to be executed once during pre-fork initialization on Snowflak
udf_init_once decorator in snowflake.snowpark.functions for marking functions to be executed once during pre-fork initialization on Snowflake workers, matching the server-side _snowflake.udf_init_once API.INFER_SCHEMA (used by DataFrameReader.csv/json/parquet/orc/avro) and COPY FILES (used by FileOperation.copy_files).COPY INTO / PUT / GET SQL, which could produce malformed statements. This affects DataFrame.write.csv/copy_into_location and the Snowpark-pandas DataFrame.to_csv stage path.SELECT query for DataFrameReader.dbapi, which could produce malformed SQL. Embedded quote characters in identifiers are now doubled (backticks for Databricks/MySQL, double quotes for Oracle/PostgreSQL/SQL Server).DataFrameWriter.copy_into_location (and csv/json/parquet/save) was embedded into the generated COPY INTO statement without quoting, which could produce malformed SQL for locations containing single quotes. The location is now consistently quoted and escaped, and a string that merely starts and ends with a single quote but contains unescaped interior quotes is no longer treated as an already-quoted literal; it is fully escaped so it stays a single SQL string literal.register_from_file were evaluated with eval(); they are now evaluated only against the documented set of supported default-value types, and unsupported expressions are ignored.object_name, object_domain, or object_version values containing single quotes or backslashes in session.lineage.trace() caused incorrect SQL generation. These values are now properly escaped before being embedded in the SYSTEM$DGQL call.comment (create_or_replace_view / dynamic table / save_as_table), collation specs (Column.collate), VARIANT/OBJECT subfield keys (Column[...]), and DataFrame/Session.flatten paths were not correctly escaped when generating SQL, which could produce malformed statements. Backslash sequences (e.g. \t, \n) in these values are now applied literally rather than interpreted.ai_extract, ai_classify, ai_similarity, ai_parse_document, ai_transcribe, ai_complete) configuration and response_format were not correctly escaped when generating the SQL object literal, which could produce malformed statements when a value contained single quotes or backslashes (for example, an apostrophe in a natural-language question).protobuf restriction for Python 3.14 from ==5.29.3 to >=5.29.3,<6.34.pandas to <3.0.0 for the [pandas] install extra, as Snowpark Python pandas-related features may not be fully compatible with pandas 3.0 or later.Added get_wif_token to snowflake.snowpark.secrets for workload identity federation tokens on the Snowflake server (not available in SPCS file-based se
get_wif_token to snowflake.snowpark.secrets for workload identity federation tokens on the Snowflake server (not available in SPCS file-based secret environments).DataFrame via copy.copy() lost post-aggregate state, causing subsequent .limit() or .sort() to generate incorrect SQL.DataFrame.alias() twice on the same DataFrame (e.g. for a self-join) caused both aliases to share the same internal column-mapping dictionary. This made col("R", "col") resolve to the same column as col("L", "col"), producing incorrect join conditions and filter expressions.cloudpickle could not be resolved when registering a Python stored procedure or UDF with runtime_version='3.13'.Clarified that the JDBC driver JAR referenced via udtf_configs.imports in DataFrameReader.jdbc() must be downloaded from the database vendor and uploa
udtf_configs.imports in DataFrameReader.jdbc() must be downloaded from the database vendor and uploaded to a Snowflake stage.Fixed a bug where AsyncJob.result("no_result") sometimes silently returned without raising error for failed queries.
AsyncJob.result("no_result") sometimes silently returned without raising error for failed queries.INTERVAL YEAR TO MONTH values with zero months were displayed incorrectly (e.g. INTERVAL '0-4' YEAR TO MONTH instead of INTERVAL '4-0' YEAR TO MONTH) when using snowflake-connector-python>=4.3.0.Session.reduce_describe_query_enabled is enabled, fewer DESCRIBE queries are issued when the outer query only projects or renames columns from an inner subquery whose column types are already known.Fixed a bug where using parameter bindings for CALL queries issued through session.sql would raise an error.
CALL queries issued through session.sql would raise an error.StringType columns from Iceberg tables were not recognized as max-size strings.FileNotFoundError message raised when INFER_SCHEMA returns zero rows so it also points to file format options (PARSE_HEADER, SKIP_HEADER, ON_ERROR=CONTINUE) that can silently filter everything out, instead of only suggesting a missing path.Added artifact_repository support to udtf_configs in session.read.dbapi(), enabling users to specify a custom artifact repository (e.g. PyPI) for pack
artifact_repository support to udtf_configs in session.read.dbapi(), enabling users to specify a custom artifact repository (e.g. PyPI) for packages used by the internal UDTF during distributed ingestion.TRY_CAST reader option is ignored when calling DataFrameReader.schema().csv().event_table_telemetry is using incorrect resource attributes.value_contains_null was not propagated for MapType in _as_nested function.uuid_string()).pandas versions to <=2.4 (was previously <=2.3.1).Allow a user-specified schema when reading Parquet files from a stage.
DataFrame.join operations.SELECT "A" AS "A" is now always simplified to SELECT "A").DataFrameReader.dbapi feature was in private preview.Session.create_dataframe raised TypeError when a StringType column was given a non-string Python value (e.g. int, float, bool, Decimal) for a small local relation (below the array bind threshold); VALUES SQL generation now coerces these types to string literals, consistent with the large-data bind-parameter path.DataFrame.approxQuantile did not accept a Column for the col parameter.concat produced extra NaN rows and mismatched values after a filter operation in local testing.dense_rank() would fail with ValueError on NULL partition values.collect() raised KeyError after save_as_table(column_order="name") when the source DataFrame omitted columns with types VariantType, MapType, or ArrayType that were present in the target table schema in local testing.Fixed a bug where chained DataFrame.filter() calls with raw SQL text containing OR produced incorrect results.
DataFrame.filter() calls with raw SQL text containing OR produced incorrect results.concat(lit('"'), ...) could lose the leading quote through chained set-operation plans.DataFrame.filter() calls with raw SQL text containing OR produced incorrect results.concat(lit('"'), ...) could lose the leading quote through chained set-operation plans.Added support for DIRECTED JOIN.
INCLUDE_METADATA copy option in DataFrame.copy_into_table, allowing users to include file metadata columns in the target table.Session.client_telemetry that trace does not have snowflake style trace id.StringType is saved in wrong length.ai_complete where model_parameters and response_format values containing single quotes would generate malformed SQL.DataFrameReader.xml() where reading XML with a custom schema whose field names contain colons (e.g., px:name) raised a SnowparkColumnException.Session.read.json when INFER_SCHEMA was set to True, and the USE_RELAXED_TYPES field of INFER_SCHEMA_OPTIONS was also set to True.SET command to Streamlit's st.write method would raise an exception.Added support for the array_union_agg function in the snowflake.snowpark.functions module.
array_union_agg function in the snowflake.snowpark.functions module.Session.udf.register_from_file did not properly process the strict and secure parameters.array_union_agg function in the snowflake.snowpark.functions module.Session.udf.register_from_file did not properly process the strict and secure parameters.DataFrame.join operations.SELECT "A" AS "A" is now always simplified to SELECT "A").Added support for the DECFLOAT data type that allows users to represent decimal numbers exactly with 38 digits of precision and a dynamic base-10 expo
DECFLOAT data type that allows users to represent decimal numbers exactly with 38 digits of precision and a dynamic base-10 exponent.DEFAULT_PYTHON_ARTIFACT_REPOSITORY parameter that allows users to configure the default artifact repository at the account, database, and schema level.cloudpickle was not automatically added to the package list when using artifact_repository with custom packages, causing ModuleNotFoundError at runtime.StructType type.Session.udf.register_from_file did not properly process the strict and secure parameters.DataFrame.join operations.SELECT "A" AS "A" is now always simplified to SELECT "A").DECFLOAT data type that allows users to represent decimal numbers exactly with 38 digits of precision and a dynamic base-10 exponent.DEFAULT_PYTHON_ARTIFACT_REPOSITORY parameter that allows users to configure the default artifact repository at the account, database, and schema level.cloudpickle was not automatically added to the package list when using artifact_repository with custom packages, causing ModuleNotFoundError at runtime.StructType type.INFER_SCHEMA with PATTERN silently fell back to unfiltered inference when the stage location had no trailing slash, causing metadata files (e.g., _common_metadata) to corrupt type inference for timestamp columns.DataFrame.join operations.SELECT "A" AS "A" is now always simplified to SELECT "A").Allow user input schema when reading XML file on stage.
functions.py:
hex_decode_stringjarowinkler_similarityparse_urlregexp_instrregexp_likeregexp_substrregexp_substr_allrtrimmed_lengthspacesplit_partpreserve_parameter_names flag to sproc, UDF, UDTF, and UDAF creationSession.client_telemetry.enable_event_table_telemetry_collection.snowflake.snowpark.context.configure_development_features is effective for multiple sessions including newly created sessions after the configuration. No duplicate experimental warning any more.DataFrame.to_arrow and DataFrame.to_arrow_batches.Session.reduce_describe_query_enabled and Session.cte_optimization_enabled are enabled, fewer DESCRIBE queries are issued when resolving table attributes.Added support for targeted delete-insert via the overwrite_condition parameter in DataFrameWriter.save_as_table
overwrite_condition parameter in DataFrameWriter.save_as_tableDataFrameReader to return columns in deterministic order when using INFER_SCHEMA.protobuf<6.34 (was <6.32).Added support for DataFrame.lateral_join
DataFrame.lateral_joinSession.client_telemetry.Session.udf_profiler.functions.ai_translate.iceberg_config options in DataFrameWriter.save_as_table and DataFrame.copy_into_table:
target_file_sizepartition_byfunctions.py:
String and Binary functions:
base64_decode_binarybucketcompressdaydecompress_binarydecompress_stringmd5_binarymd5_number_lower64md5_number_upper64sha1_binarysha2_binarysoundex_p123strtoktruncatetry_base64_decode_binarytry_base64_decode_stringtry_hex_decode_binarytry_hex_decode_stringunicodeuuid_stringConditional expressions:
booland_aggboolxor_aggregr_valyzeroifnullNumeric expressions:
cotmodpisquarewidth_bucketDataFrames created using DataFrame.alias and CTE optimization is enabled.XMLReader where finding the start position of a row tag could return an incorrect file position.DataFrame.sort() to support ORDER BY ALL when no columns are specified.Session.cte_optimization_enabled.Dataframe.groupby.rolling().np.percentile with DataFrame and Series inputs to Series.quantile.random_state parameter to an integer when calling DataFrame.sample or Series.sample.iceberg_config options in to_iceberg:
target_file_sizepartition_byshift() with suffix or non-integer periods parameterssort_index() with axis=1 or key parameterssort_values() with axis=1melt() with col_level parameterapply() with result_type parameter for DataFramepivot_table() with sort=True, non-string index list, non-string columns list, non-string values list, or aggfunc dict with non-string valuesfillna() with downcast parameter or using limit together with valuedropna() with axis=1asfreq() with how parameter, fill_value parameter, normalize=True, or freq parameter being week, month, quarter, or yeargroupby() with axis=1, by!=None and level!=None, or by containing any non-pandas hashable labels.groupby_fillna() with downcast parametergroupby_first() with min_count>1groupby_last() with min_count>1groupby_shift() with freq parameteragg, nunique, describe, and related methods on 1-column DataFrame and Series objects.DataFrameGroupBy.agg where func is a list of tuples used to set the names of the output columns.np.asarray would cause a TypeError.Series.isin with a Series argument matched index labels instead of the row position.groupby.applygroupby.nuniquegroupby.sizeconcatcopystr.isdigitstr.islowerstr.isupperstr.istitlestr.lowerstr.upperstr.titlestr.matchstr.capitalizestr.__getitem__str.centerstr.countstr.getstr.padstr.lenstr.ljuststr.rjuststr.splitstr.replacestr.stripstr.lstripstr.rstripstr.translatedt.tz_localizedt.tz_convertdt.ceildt.rounddt.floordt.normalizedt.month_namedt.day_namedt.strftimedt.dayofweekdt.weekdaydt.dayofyeardt.isocalendarrolling.minrolling.maxrolling.countrolling.sumrolling.meanrolling.stdrolling.varrolling.semrolling.correxpanding.minexpanding.maxexpanding.countexpanding.sumexpanding.meanexpanding.stdexpanding.varexpanding.semcumsumcummincummaxgroupby.groupsgroupby.indicesgroupby.firstgroupby.lastgroupby.rankgroupby.shiftgroupby.cumcountgroupby.cumsumgroupby.cummingroupby.cummaxgroupby.anygroupby.allgroupby.uniquegroupby.get_groupgroupby.rollinggroupby.resampleto_snowflaketo_snowparkresample.minresample.maxresample.countresample.sumresample.meanresample.medianresample.stdresample.varresample.sizeresample.firstresample.lastresample.quantileresample.nuniquedrop_duplicates by avoiding joins when keep!=False in faster pandas.Snowpark python DB-api is now generally available. Access this feature with DataFrameReader.dbapi() to read data from a database table or query into a
DataFrameReader.dbapi() to read data from a database table or query into a DataFrame using a DBAPI connection.Added a new function service in snowflake.snowpark.functions that allows users to create a callable representing a Snowpark Container Services (SPCS)
service in snowflake.snowpark.functions that allows users to create a callable representing a Snowpark Container Services (SPCS) service.connection_parameters parameter to DataFrameReader.dbapi() (PuPr) method to allow passing keyword arguments to the create_connection callable.Session.begin_transaction, Session.commit and Session.rollback.functions.py:
st_interpolatest_intersectionst_intersection_aggst_intersectsst_isvalidst_lengthst_makegeompointst_makelinest_makepolygonst_makepolygonorientedst_disjointst_distancest_dwithinst_endpointst_envelopest_geohashst_geomfromgeohashst_geompointfromgeohashst_hausdorffdistancest_makepointst_npointsst_perimeterst_pointnst_setsridst_simplifyst_sridst_startpointst_symdifferencest_transformst_unionst_union_aggst_withinst_xst_xmaxst_xminst_yst_ymaxst_yminst_geogfromgeohashst_geogpointfromgeohashst_geographyfromwkbst_geographyfromwktst_geometryfromwkbst_geometryfromwkttry_to_geographytry_to_geometryinterval_day_time_from_parts and interval_year_month_from_parts functions.DataFrameReader.xml fails to parse XML files with undeclared namespaces when ignoreNamespace is True.interval_day_time_from_parts.to_snowflake would raise KeyError.DataFrameReader.dbapi (PuPr) is not compatible with oracledb 3.4.0.modin would unintentionally be imported during session initialization in some scenarios.session.udf|udtf|udaf|sproc.register failed when an extra session argument was passed. These methods do not expect a session argument; please remove it if provided.DataFrameReader.dbapi is now increased from 16MB to 128MB in parquet file based ingestion.snowflake-connector-python>=3.17,<5.0.0.dtypes parameter of pd.get_dummiesnunique in df.pivot_table, df.agg and other places where aggregate functions can be used.DataFrame.interpolate and Series.interpolate with the "linear", "ffill"/"pad", and "backfill"/bfill" methods. These use the SQL INTERPOLATE_LINEAR, INTERPOLATE_FFILL, and INTERPOLATE_BFILL functions (PuPr).Series.to_snowflake and pd.to_snowflake(series) for large data by uploading data via a parquet file. You can control the dataset size at which Snowpark pandas switches to parquet with the variable modin.config.PandasToSnowflakeParquetThresholdBytes.get_dummies() with dummy_na=True, drop_first=True, or custom dtype parameterscumsum(), cummin(), cummax() with axis=1 (column-wise operations)skew() with axis=1 or numeric_only=False parametersround() with decimals parameter as a Seriescorr() with method!=pearson parametercte_optimization_enabled to True for all Snowpark pandas sessions.isinisnaisnullnotnanotnullstr.containsstr.startswithstr.endswithstr.slicedt.datedt.timedt.hourdt.minutedt.seconddt.microseconddt.nanoseconddt.yeardt.monthdt.daydt.quarterdt.is_month_startdt.is_month_enddt.is_quarter_startdt.is_quarter_enddt.is_year_startdt.is_year_enddt.is_leap_yeardt.days_in_monthdt.daysinmonthsort_valuesloc (setting columns)to_datetimerenamedropinvertduplicatedilocheadcolumns (e.g., df.columns = ["A", "B"])aggminmaxcountsummeanmedianstdvargroupby.agggroupby.mingroupby.maxgroupby.countgroupby.sumgroupby.meangroupby.mediangroupby.stdgroupby.vardrop_duplicatesget_axis_len.Added a new module snowflake.snowpark.secrets that provides Python wrappers for accessing Snowflake Secrets within Python UDFs and stored procedures t
Added a new module snowflake.snowpark.secrets that provides Python wrappers for accessing Snowflake Secrets within Python UDFs and stored procedures that execute inside Snowflake.
get_generic_secret_stringget_oauth_access_tokenget_secret_typeget_username_passwordget_cloud_provider_tokenAdded support for the following scalar functions in functions.py:
Conditional expression functions:
boolandboolnotboolorboolxorboolor_aggdecodegreatest_ignore_nullsleast_ignore_nullsnullifnvl2regr_valxSemi-structured and structured date functions:
array_remove_atas_booleanmap_deletemap_insertmap_pickmap_sizeString & binary functions:
chrhex_decode_binaryNumeric functions:
div0nullDifferential privacy functions:
dp_interval_highdp_interval_lowContext functions:
last_query_idlast_transactionGeospatial functions:
h3_cell_to_boundaryh3_cell_to_childrenh3_cell_to_children_stringh3_cell_to_parenth3_cell_to_pointh3_compact_cellsh3_compact_cells_stringsh3_coverageh3_coverage_stringsh3_get_resolutionh3_grid_diskh3_grid_distanceh3_int_to_stringh3_polygon_to_cellsh3_polygon_to_cells_stringsh3_string_to_inth3_try_grid_pathh3_try_polygon_to_cellsh3_try_polygon_to_cells_stringsh3_uncompact_cellsh3_uncompact_cells_stringshaversineh3_grid_pathh3_is_pentagonh3_is_valid_cellh3_latlng_to_cellh3_latlng_to_cell_stringh3_point_to_cellh3_point_to_cell_stringh3_try_coverageh3_try_coverage_stringsh3_try_grid_distancest_areast_asewkbst_asewktst_asgeojsonst_aswkbst_aswktst_azimuthst_bufferst_centroidst_collectst_containsst_coveredbyst_coversst_differencest_dimensionDataFrame.limit() fail if there is parameter binding in the executed SQL when used in non-stored-procedure/udxf environment.DataFrameReader.dbapi (PuPr):
pyodbc driver caused by unprocessed row data.Session.create_dataframe where schema string parsing failed when using upper-cased data type names (e.g., NUMERIC, NUMBER, DECIMAL, VARCHAR, STRING, TEXT).DataFrameReader.dbapi(PuPr) that dbapi will not retry on non-retryable error such as SQL syntax error on external data source query.session.read.option('rowTag', <tag_name>).xml(<stage_file_path>) or xpath functions.DataFrameReader.dbapi (PuPr) reading performance by setting the default fetch_size parameter value to 100000.session.read.option('rowValidationXSDPath', <xsd_path>).xml(<stage_file_path>).modin versions to >=0.36.0 and <0.38.0 (was previously >= 0.35.0 and <0.37.0).DataFrame.query for dataframes with single-level indexes.DataFrameGroupby.__len__ and SeriesGroupBy.__len__.from modin.config import AutoSwitchBackend; AutoSwitchBackend.disable() to turn this off and force all execution to occur in Snowflake.pandas_hybrid_execution_enabled to enable/disable hybrid execution as an alternative to using AutoSwitchBackend.SHOW OBJECTS query issued from read_snowflake under certain conditions.pd.merge, pd.concat, DataFrame.merge, and DataFrame.join may now move arguments to backends other than those among the function arguments.DataFrame.to_snowflake and pd.to_snowflake(dataframe) for large data by uploading data via a parquet file. You can control the dataset size at which Snowpark pandas switches to parquet with the variable modin.config.PandasToSnowflakeParquetThresholdBytes.Added an experimental fix for a bug in schema query generation that could cause invalid sql to be genrated when using nested structured types.
Deprecated warnings will be triggered when using snowpark-python with Python 3.9. For more details, please refer to https://docs.snowflake.com/en/deve…
DataFrame.ai.complete: Generate per-row LLM completions from prompts built over columns and files.DataFrame.ai.filter: Keep rows where an AI classifier returns TRUE for the given predicate.DataFrame.ai.agg: Reduce a text column into one result using a natural-language task description.RelationalGroupedDataFrame.ai_agg: Perform the same natural-language aggregation per group.DataFrame.ai.classify: Assign single or multiple labels from given categories to text or images.DataFrame.ai.similarity: Compute cosine-based similarity scores between two columns via embeddings.DataFrame.ai.sentiment: Extract overall and aspect-level sentiment from text into JSON.DataFrame.ai.embed: Generate VECTOR embeddings for text or images using configurable models.DataFrame.ai.summarize_agg: Aggregate and produce a single comprehensive summary over many rows.DataFrame.ai.transcribe: Transcribe audio files to text with optional timestamps and speaker labels.DataFrame.ai.parse_document: OCR/layout-parse documents or images into structured JSON.DataFrame.ai.extract: Pull structured fields from text or files using a response schema.DataFrame.ai.count_tokens: Estimate token usage for a given model and input text per row.DataFrame.ai.split_text_markdown_header: Split Markdown into hierarchical header-aware chunks.DataFrame.ai.split_text_recursive_character: Split text into size-bounded chunks using recursive separators.DataFrameReader.file: Create a DataFrame containing all files from a stage as FILE data type for downstream unstructured data processing.YearMonthIntervalType that allows users to create intervals for datetime operations.interval_year_month_from_parts that allows users to easily create YearMonthIntervalType without using SQL.DayTimeIntervalType that allows users to create intervals for datetime operations.interval_day_time_from_parts that allows users to easily create DayTimeIntervalType without using SQL.FileOperation.list to list files in a stage with metadata.FileOperation.remove to remove files in a stage.copy_grants for the following DataFrame APIs:
create_or_replace_viewcreate_or_replace_temp_viewcreate_or_replace_dynamic_tablesnowflake.snowpark.functions.vectorized that allows users to mark a function as vectorized UDF.use_vectorized_scanner in function Session.write_pandas().functions.py:
getdategetvariableinvoker_roleinvoker_shareis_application_role_in_sessionis_database_role_in_sessionis_granted_to_invoker_roleis_role_in_sessionlocaltimesystimestampDataFrameReader.dbapi(PuPr) are ingested as StringType now.cacheResult to DataFrameReader.xml that allows users to cache the result of the XML reader to a temporary table after calling xml. It helps improve performance when subsequent operations are performed on the same DataFrame.logging.DEBUG - 1 the log message saying that the
Snowpark DataFrame reference of an internal DataFrameReference object
has changed.Complete.read_snowflake, repr, loc, reset_index, merge, and binary operations.dummy_row_pos_optimization_enabled to enable/disable dummy row position optimization in faster pandas.modin versions to >=0.35.0 and <0.37.0 (was previously >= 0.34.0 and <0.36.0).AssertionError was unexpectedly raised by certain indexing operations.functions.ai_complete.Added support for the following AI-powered functions in functions.py:
functions.py:
ai_extractai_parse_documentai_transcribeSession.table() now supports time travel parameters: time_travel_mode, statement, offset, timestamp, timestamp_type, and stream.DataFrameReader.table() supports the same time travel parameters as direct arguments.DataFrameReader supports time travel via option chaining (e.g., session.read.option("time_travel_mode", "at").option("offset", -60).table("my_table")).DataFrameWriter.copy_into_location for validation and writing data to external locations:
validation_modestorage_integrationcredentialsencryptionSession.directory and Session.read.directory to retrieve the list of all files on a stage with metadata.DataFrameReader.jdbc(PrPr) that allows ingesting external data source with jdbc driver.FileOperation.copy_files to copy files from a source location to an output stage.functions.py:
all_user_namesbitandbitand_aggbitorbitor_aggbitxorbitxor_aggcurrent_account_namecurrent_clientcurrent_ip_addresscurrent_role_typecurrent_organization_namecurrent_organization_usercurrent_secondary_rolescurrent_transactiongetbitDataFrameReader.dbapi that udtf ingestion does not work in stored procedure.DataFrameReader.dbapi thread-based ingestion to prevent unnecessary operations, which improves resource efficiency.cloudpickle==3.1.1 in addition to previous versions.DataFrameReader.dbapi (PuPr) ingestion performance for PostgreSQL and MySQL by using server side cursor to fetch data.pd.read_snowflake(), pd.to_iceberg(),
pd.to_pandas(), pd.to_snowpark(), pd.to_snowflake(),
DataFrame.to_iceberg(), DataFrame.to_pandas(), DataFrame.to_snowpark(),
DataFrame.to_snowflake(), Series.to_iceberg(), Series.to_pandas(),
Series.to_snowpark(), and Series.to_snowflake() on the "Pandas" and "Ray"
backends. Previously, only some of these functions and methods were supported
on the Pandas backend.Index.get_level_values().--upgrade to pip install "snowflake-snowpark-python[modin]" in the error message.NotImplementedError instead of AttributeError on attempting to call
Snowflake extension functions/methods to_dynamic_table(), cache_result(),
to_view(), create_or_replace_dynamic_table(), and
create_or_replace_view() on dataframes or series using the pandas or ray
backends.Added support for the following xpath functions in functions.py:
xpath functions in functions.py:
xpathxpath_stringxpath_booleanxpath_intxpath_floatxpath_doublexpath_longxpath_shortuse_vectorized_scanner in function Session.write_arrow().DataFrame.col_ilike.AsyncJob objects.
block: bool = True parameter to Session.call(). When block=False, returns an AsyncJob instead of blocking until completion.block: bool = True parameter to StoredProcedure.__call__() for async support across both named and anonymous stored procedures.Session.call_nowait() that is equivalent to Session.call(block=False).deepcopy of internal plans would cause a memory spike when a dataframe is created locally using session.create_dataframe() using a large input data.DataFrameReader.parquet where the ignore_case option in the infer_schema_options was not respected.to_pandas() has different format of column name when query result format is set to 'JSON' and 'ARROW'.pkg_resources.protobuf<6.32DataFrame.set_backend method. The installed version of modin must be at least 0.35.0, and ray must be installed.modin versions to >=0.34.0 and <0.36.0 (was previously >= 0.33.0 and <0.35.0).modin version is at least 0.35.0.pd.to_datetime and pd.to_timedelta would unexpectedly raise IndexError.pd.explain_switch would raise IndexError or return None if called before any potential switch operations were performed.Session.create_dataframe now accepts keyword arguments that are forwarded to the internal call to Session.write_pandas or Session.write_arrow when cre
Session.create_dataframe now accepts keyword arguments that are forwarded to the internal call to Session.write_pandas or Session.write_arrow when creating a DataFrame from a pandas DataFrame or a pyarrow Table.AsyncJob:
AsyncJob.is_failed() returns a bool indicating if a job has failed. Can be used in combination with AsyncJob.is_done() to determine if a job is finished and errored.AsyncJob.status() returns a string representing the current query status (e.g., "RUNNING", "SUCCESS", "FAILED_WITH_ERROR") for detailed monitoring without calling result().functions.py:
ai_sentimentcontext.configure_development_features. All development features are disabled by default unless explicitly enabled by the user.DataFrame/Series/GroupBy.apply, map, and transform by passing the snowflake_udf_params keyword argument. See documentation for details.AutoSwitchBackend even when users had explicitly configured it via environment variables or programmatically.Added support for the following functions in functions.py:
functions.py:
ai_embedtry_parse_jsonDataFrameReader.dbapi (PrPr) that dbapi fail in python stored procedure with process exit with code 1.DataFrameReader.dbapi (PrPr) that custom_schema accept illegal schema.DataFrameReader.dbapi (PrPr) that custom_schema does not work when connecting to Postgres and Mysql.query parameter in DataFrameReader.dbapi (PrPr) so that parentheses are not needed around the query.DataFrameReader.dbapi (PrPr) when exception happen during inferring schema of target data source.SnowflakeFile using local file paths, the Snow URL semantic (snow://...), local testing framework stages, and Snowflake stages (@stage/file_path).DataFrame.boxplot.apply or map with the same arguments on Snowpark pandas objects.pd.read_excel bug when reading files inside stage inner directory.functions.py:
ai_embedtry_parse_jsonDataFrameReader.dbapi (PrPr) that dbapi fail in python stored procedure with process exit with code 1.DataFrameReader.dbapi (PrPr) that custom_schema accept illegal schema.DataFrameReader.dbapi (PrPr) that custom_schema does not work when connecting to Postgres and Mysql.query parameter in DataFrameReader.dbapi (PrPr) so that parentheses are not needed around the query.DataFrameReader.dbapi (PrPr) when exception happen during inferring schema of target data source.SnowflakeFile using local file paths, the Snow URL semantic (snow://...), local testing framework stages, and Snowflake stages (@stage/file_path).DataFrame.boxplot.apply or map with the same arguments on Snowpark pandas objects.pd.read_excel.pd.read_excel bug when reading files inside stage inner directory.Added a new option TRY_CAST to DataFrameReader. When TRY_CAST is True columns are wrapped in a TRY_CAST statement rather than a hard cast when loading
TRY_CAST to DataFrameReader. When TRY_CAST is True columns are wrapped in a TRY_CAST statement rather than a hard cast when loading data.USE_RELAXED_TYPES to the INFER_SCHEMA_OPTIONS of DataFrameReader. When set to True this option casts all strings to max length strings and all numeric types to DoubleType.snowflake.snowpark.context.configure_development_features().snowflake.snowpark.dataframe.map_in_pandas that allows users map a function across a dataframe. The mapping function takes an iterator of pandas dataframes as input and provides one as output.fetch_with_process to DataFrameReader.dbapi (PrPr) to enable multiprocessing for parallel data fetching in
local ingestion. By default, local ingestion uses multithreading. Multiprocessing may improve performance for CPU-bound tasks like Parquet file generation.snowflake.snowpark.functions.model that allows users to call methods of a model.rowValidationXSDPath option when reading XML files with a row tag using rowTag option.session.table().sample() to generate a flat SQL statement.functions.explode.snowflake.snowpark.context.configure_development_features(). This feature also depends on AST collection to be enabled in the session which can be done using session.ast_enabled = True.to_snowpark_pandas() from a snowpark dataframe containing DML/DDL queries instead of throwing a NotImplementedError.DataFrameReader.dbapi (PrPr) where closing the cursor or connection could unexpectedly raise an error and terminate the program.DataFrame.select() that have output columns matching the input DataFrame's columns. This improvement works when dataframe columns are provided as Column objects.DataFrame.to_excel and Series.to_excel.pd.read_feather, pd.read_orc, and pd.read_stata.pd.explain_switch() to return debugging information on hybrid execution decisions.pd.read_snowflake when the global modin backend is Pandas.pd.to_dynamic_table, pd.to_iceberg, and pd.to_view.modin or pandas version does not match our requirements.modin versions to >=0.33.0 and <0.35.0 (was previously >= 0.32.0 and <0.34.0).TypeError: numpy.ndarray object is not callable.np.where on modin objects with the Pandas backend would raise an AttributeError. This fix requires modin version 0.34.0 or newer.df.melt where the resulting values have an additional suffix applied.Added support for MySQL in DataFrameWriter.dbapi (PrPr) for both Parquet and UDTF-based ingestion.
DataFrameWriter.dbapi (PrPr) for both Parquet and UDTF-based ingestion.DataFrameReader.dbapi (PrPr) for both Parquet and UDTF-based ingestion.DataFrameWriter.dbapi (PrPr) for UDTF-based ingestion.DataFrameReader to enable use of PATTERN when reading files with INFER_SCHEMA enabled.functions.py:
ai_completeai_similarityai_summarize_agg (originally summarize_agg)ai_classifyrowTag option:
ignoreNamespace option.attributePrefix option.excludeAttributes option.valueTag option.null value using nullValue option.charset option.ignoreSurroundingWhitespace option.return_dataframe in Session.call, which can be used to set the return type of the functions to a DataFrame object.Dataframe.describe called strings_include_math_stats that triggers stddev and mean to be calculated for String columns.Edge.properties when retrieving lineage from DGQL in DataFrame.lineage.trace.table_exists to DataFrameWriter.save_as_table that allows specifying if a table already exists. This allows skipping a table lookup that can be expensive.DataFrameReader.dbapi (PrPr) where the create_connection defined as local function was incompatible with multiprocessing.DataFrameReader.dbapi (PrPr) where databricks TIMESTAMP type was converted to Snowflake TIMESTAMP_NTZ type which should be TIMESTAMP_LTZ type.DataFrameReader.json where repeated reads with the same reader object would create incorrectly quoted columns.DataFrame.to_pandas() that would drop column names when converting a dataframe that did not originate from a select statement.DataFrame.create_or_replace_dynamic_table raises error when the dataframe contains a UDTF and SELECT * in UDTF not being parsed correctly.Session.write_pandas() and Session.create_dataframe() when the input pandas DataFrame does not have a column.DataFrame.select when the arguments contain a table function with output columns that collide with columns of current dataframe. With the improvement, if user provides non-colliding columns in df.select("col1", "col2", table_func(...)) as string arguments, then the query generated by snowpark client will not raise ambiguous column error.DataFrameReader.dbapi (PrPr) to use in-memory Parquet-based ingestion for better performance and security.DataFrameReader.dbapi (PrPr) to use MATCH_BY_COLUMN_NAME=CASE_SENSITIVE in copy into table operation.Column.isin that would cause incorrect filtering on joined or previously filtered data.snowflake.snowpark.functions.concat_ws that would cause results to have an incorrect index.modin dependency constraint from 0.32.0 to >=0.32.0, <0.34.0. The latest version tested with Snowpark pandas is modin 0.33.1.from modin.config import AutoSwitchBackend; AutoSwitchBackend.enable(), Snowpark pandas will automatically choose whether to run certain pandas operations locally or on Snowflake. This feature is disabled by default.index parameter to False for DataFrame.to_view, Series.to_view, DataFrame.to_dynamic_table, and Series.to_dynamic_table.iceberg_version option to table creation functions.insert, repr, and groupby, that previously issued a query to retrieve the input data's size.Series.where when the other parameter is an unnamed Series.Invoking snowflake system procedures does not invoke an additional describe procedure call to check the return type of the procedure.
describe procedure call to check the return type of the procedure.Session.create_dataframe() with the stage URL and FILE data type.session.read.option('mode', <mode>), option('rowTag', <tag_name>).xml(<stage_file_path>). Currently PERMISSIVE, DROPMALFORMED and FAILFAST are supported.Dataframe.drop to use SELECT * EXCLUDE () to exclude the dropped columns. To enable this feature, set session.conf.set("use_simplified_query_generation", True).VariantType to StructType.from_jsonDataFrameWriter.dbapi (PrPr) that unicode or double-quoted column name in external database causes error because not quoted correctly.snowflake.snowpark.functions.rank that would cause sort direction to not be respected.snowflake.snowpark.functions.to_timestamp_* that would cause incorrect results on filtered data.Series.str.get, Series.str.slice, and Series.str.__getitem__ (Series.str[...]).DataFrame.to_html.DataFrame.to_string and Series.to_string.pd.read_csv.iceberg_config a required parameter for DataFrame.to_iceberg and Series.to_iceberg.describe procedure call to check the return type of the procedure.Session.create_dataframe() with the stage URL and FILE data type.session.read.option('mode', <mode>), option('rowTag', <tag_name>).xml(<stage_file_path>). Currently PERMISSIVE, DROPMALFORMED and FAILFAST are supported.Dataframe.drop to use SELECT * EXCLUDE () to exclude the dropped columns. To enable this feature, set session.conf.set("use_simplified_query_generation", True).VariantType to StructType.from_jsonDataFrameWriter.dbapi (PrPr) that unicode or double-quoted column name in external database causes error because not quoted correctly.native_app_params parameters in register udaf function.snowflake.snowpark.functions.rank that would cause sort direction to not be respected.snowflake.snowpark.functions.to_timestamp_* that would cause incorrect results on filtered data.Series.str.get, Series.str.slice, and Series.str.__getitem__ (Series.str[...]).DataFrame.to_html.DataFrame.to_string and Series.to_string.pd.read_csv.ENFORCE_EXISTING_FILE_FORMAT option to the DataFrameReader, which allows to read a dataframe only based on an existing file format object when used together with FORMAT_NAME.iceberg_config a required parameter for DataFrame.to_iceberg and Series.to_iceberg.Updated conda build configuration to deprecate Python 3.8 support, preventing installation in incompatible environments.
Deprecated support for Python3.8.
restricted caller permission of execute_as argument in StoredProcedure.register().DataFrame.to_pandas().artifact_repository parameter to Session.add_packages, Session.add_requirements, Session.get_packages, Session.remove_package, and Session.clear_packages.session.read.option('rowTag', <tag_name>).xml(<stage_file_path>) (experimental).
col(a.b.c).DataFrameReader.dbapi (PrPr):
fetch_merge_count parameter for optimizing performance by merging multiple fetched data into a single Parquet file.functions.py (Private Preview):
promptai_filter (added support for prompt() function and image files, and changed the second argument name from expr to file)ai_classifyrelaxed_ordering param into enforce_ordering for DataFrame.to_snowpark_pandas. Also the new default values is enforce_ordering=False which has the opposite effect of the previous default value, relaxed_ordering=False.DataFrameReader.dbapi (PrPr) reading performance by setting the default fetch_size parameter value to 1000.session.table.DataFrameAnalyticsFunctions.time_series_agg().DataFrame.group_by().pivot().agg when the pivot column and aggregate column are the same.DataFrameReader.dbapi (PrPr) where a TypeError was raised when create_connection returned a connection object of an unsupported driver type.df.limit(0) call would not properly apply.DataFrameWriter.save_as_table that caused reserved names to throw errors when using append mode.sliding_interval in DataFrameAnalyticsFunctions.time_series_agg().Window.range_between.array_construct function.__pycache__ directory was unintentionally copied during stored procedure execution via import.Column.like calls.Column.getItem and snowpark.snowflake.functions.get to raise IndexError rather than return null.df.limit(0) call would not properly apply.Table.merge into an empty table would cause an exception.modin from 0.30.1 to 0.32.0.numpy 2.0 and above.DataFrame.create_or_replace_view and Series.create_or_replace_view.DataFrame.create_or_replace_dynamic_table and Series.create_or_replace_dynamic_table.DataFrame.to_view and Series.to_view.DataFrame.to_dynamic_table and Series.to_dynamic_table.DataFrame.groupby.resample for aggregations max, mean, median, min, and sum.pd.read_excelpd.read_htmlpd.read_picklepd.read_saspd.read_xmlDataFrame.to_iceberg and Series.to_iceberg.Series.str.len.DataFrame.groupby.apply and Series.groupby.apply by avoiding expensive pivot step.OrderedDataFrame to enable better engine switching. This could potentially result in increased query counts.relaxed_ordering param into enforce_ordering for pd.read_snowflake. Also the new default value is enforce_ordering=False which has the opposite effect of the previous default value, relaxed_ordering=False.pd.read_snowflake when reading iceberg tables and enforce_ordering=True.Added Support for relaxed consistency and ordering guarantees in Dataframe.to_snowpark_pandas by introducing the new parameter relaxed_ordering.
Dataframe.to_snowpark_pandas by introducing the new parameter relaxed_ordering.DataFrameReader.dbapi (PrPr) now accepts a list of strings for the session_init_statement parameter, allowing multiple SQL statements to be executed during session initialization.Dataframe.stat.sample_by to generate a single flat query that scales well with large fractions dictionary compared to older method of creating a UNION ALL subquery for each key in fractions. To enable this feature, set session.conf.set("use_simplified_query_generation", True).DataFrameReader.dbapi by enable vectorized option when copy parquet file into table.DataFrame.random_split in the following ways. They can be enabled by setting session.conf.set("use_simplified_query_generation", True):
cache_result in the internal implementation of the input dataframe resulting in a pure lazy dataframe operation.seed argument now behaves as expected with repeatable results across multiple calls and sessions.DataFrame.fillna and DataFrame.replace now both support fitting int and float into Decimal columns if include_decimal is set to True.files.py as a result of their General Availability.
SnowflakeFile.writeSnowflakeFile.writelinesSnowflakeFile.writeableSnowflakeFile and SnowflakeFile.open().cast() is applied to their output
from_jsonsizeDataframe.except_ that would cause rows to be incorrectly dropped.to_timestamp to fail when casting filtered columns.Series.str.__getitem__ (Series.str[...]).pd.Grouper objects in group by operations. When freq is specified, the default values of the sort, closed, label, and convention arguments are supported; origin is supported when it is start or start_day.pd.read_snowflake for both named data sources (e.g., tables and views) and query data sources by introducing the new parameter relaxed_ordering.QUOTED_IDENTIFIERS_IGNORE_CASE is found to be set, ask user to unset it.index_label in DataFrame.to_snowflake and Series.to_snowflake is handled when index=True. Instead of raising a ValueError, system-defined labels are used for the index columns.groupby or DataFrame or Series.agg when the function name is not supported.Fixed a bug in DataFrameReader.dbapi (PrPr) that prevents usage in stored procedure and snowbooks.
DataFrameReader.dbapi (PrPr) that prevents usage in stored procedure and snowbooks.Added support for the following AI-powered functions in functions.py (Private Preview):
functions.py (Private Preview):
ai_filterai_aggsummarize_aggfunctions.py (Private Preview):
fl_get_content_typefl_get_etagfl_get_file_typefl_get_last_modifiedfl_get_relative_pathfl_get_scoped_file_urlfl_get_sizefl_get_stagefl_get_stage_file_urlfl_is_audiofl_is_compressedfl_is_documentfl_is_imagefl_is_videoartifact_repository and artifact_repository_packages to specify your artifact repository and packages respectively when registering stored procedures or user defined functions.Session.sproc.registerSession.udf.registerSession.udaf.registerSession.udtf.registerfunctions.sprocfunctions.udffunctions.udaffunctions.udtffunctions.pandas_udffunctions.pandas_udtfUnsupported feature 'SCOPED_TEMPORARY'. error if thread-safe session was disabled.df.describe raised internal SQL execution error when the dataframe is created from reading a stage file and CTE optimization is enabled.df.order_by(A).select(B).distinct() would generate invalid SQL when simplified query generation was enabled using session.conf.set("use_simplified_query_generation", True).
snowflake-snowpark-python package compatibility when registering stored procedures. Now, warnings are only triggered if the major or minor version does not match, while bugfix version differences no longer generate warnings.cloudpickle==3.0.0 in addition to previous versions.range_between window function.ClassifyText, Translate, and ExtractAnswer.pd.to_snowflake, DataFrame.to_snowflake, and Series.to_snowflake when the table does not exist.if_exists parameter in pd.to_snowflake, DataFrame.to_snowflake, and Series.to_snowflake.Series.rename_axis where an AttributeError was being raised.pd.get_dummies didn't ignore NULL/NaN values by default.pd.get_dummies results in 'Duplicated column name error'.pd.get_dummies where passing list of columns generated incorrect column labels in output DataFrame.pd.get_dummies to return bool values instead of int.Deprecated Snowpark Python function snowflake_cortex_summarize. Users can install snowflake-ml-python and use the snowflake.cortex.summarize function…
functions.py
normalrandnallow_missing_columns parameter to Dataframe.union_by_name and Dataframe.union_all_by_name.Dataframe.distinct to generate SELECT DISTINCT instead of SELECT with GROUP BY all columns. To disable this feature, set session.conf.set("use_simplified_query_generation", False).snowflake_cortex_summarize. Users can install snowflake-ml-python and use the snowflake.cortex.summarize function instead.snowflake_cortex_sentiment. Users can install snowflake-ml-python and use the snowflake.cortex.sentiment function instead.session.conf.set("collect_stacktrace_in_query_tag", True).Session._write_pandas where it was erroneously passing use_logical_type parameter to Session._write_modin_pandas_helper when writing a Snowpark pandas object.Session.catalog where empty strings for database or schema were not handled correctly and were generating erroneous sql statements.Summarize and Sentiment.Series.str.get.apply where kwargs were not being correctly passed into the applied function.hourminutedate_format, datetime_format, and timestamp_format options when loading csvs.Added support for the following functions in functions.py
functions.py
array_reversedivnullmap_catmap_contains_keymap_keysnullifzerosnowflake_cortex_sentimentacoshasinhatanhbit_lengthbitmap_bit_positionbitmap_bucket_numberbitmap_construct_aggcbrtequal_nullfrom_jsonifnulllocaltimestampmax_bymin_bynth_valuenvloctet_lengthpositionregr_avgxregr_avgyregr_countregr_interceptregr_r2regr_sloperegr_sxxregr_sxyregr_syytry_to_binarybase64base64_decode_stringbase64_encodeeditdistancehexhex_encodeinstrlog1plog2log10percentile_approxunbase64DataFrame.create_dataframe.DataFrameWriter.insert_into/insertInto. This method also supports local testing mode.DataFrame.create_temp_view to create a temporary view. It will fail if the view already exists.map_cat and map_concat.keep_column_order for keeping original column order in DataFrame.with_column and DataFrame.with_columns.contains_null parameter to ArrayType.DataFrame.create_or_replace_temp_view from a DataFrame created by reading a file from a stage.value_contains_null parameter to MapType.interactive to telemetry that indicates whether the current environment is an interactive one.session.file.get in a Native App to read file paths starting with / from the current versionDataFrame.pivot.Catalog class to manage snowflake objects. It can be accessed via Session.catalog.
snowflake.core is a dependency required for this feature.DataFrame.create_dataframe.cosign.StructField.from_json that prevented TimestampTypes with tzinfo from being parsed correctly.date_format that caused an error when the input column was date type or timestamp type.replace and lit which raised type hint assertion error when passing Column expression objects.pandas_udf and pandas_udtf where session parameter was erroneously ignored.session.call.Series.str.ljust and Series.str.rjust.Series.str.center.Series.str.pad.snowflake_cortex_sentiment.DataFrame.map.DataFrame.from_dict and DataFrame.from_records.SeriesGroupBy.uniqueSeries.dt.strftime with the following directives:
Series.between.include_groups=False in DataFrameGroupBy.apply.expand=True in Series.str.split.DataFrame.pop and Series.pop.first and last in DataFrameGroupBy.agg and SeriesGroupBy.agg.Index.drop_duplicates."count", "median", np.median,
"skew", "std", np.std "var", and np.var in
pd.pivot_table(), DataFrame.pivot_table(), and pd.crosstab().DataFrame.map, Series.apply and Series.map methods by mapping numpy functions to snowpark functions if possible.DataFrame.map.DataFrame.apply by mapping numpy functions to snowpark functions if possible.Series.map, Series.apply and DataFrame.map if type-hint is not provided.call_count to telemetry that counts method calls including interchange protocol calls.Added support for property version and class method get_active_session for Session class.
version and class method get_active_session for Session class.DataType, its derived classes, and StructField:
type_name: Returns the type name of the data.simple_string: Provides a simple string representation of the data.json_value: Returns the data as a JSON-compatible value.json: Converts the data to a JSON string.ArrayType, MapType, StructField, PandasSeriesType, PandasDataFrameType and StructType:
from_json: Enables these types to be created from JSON data.MapType:
keyType: keys of the mapvalueType: values of the mapappName in SessionBuilder.include_nulls argument in DataFrame.unpivot.functions.py:
size to get size of array, object, or map columns.collect_list an alias of array_agg.substring makes len argument optional.ast_enabled to session for internal usage (default: False).DataFrame.create_or_replace_dynamic_table:
iceberg_config A dictionary that can hold the following iceberg configuration options:
external_volumecatalogbase_locationcatalog_syncstorage_serialization_policyDataFrame.print_schemalevel parameter to DataFrame.print_schemaDataFrameReader and DataFrameWriter API by adding support for the following:
format method to DataFrameReader and DataFrameWriter to specify file format when loading or unloading results.load method to DataFrameReader to work in conjunction with format.save method to DataFrameWriter to work in conjunction with format.options method for DataFrameReader and DataFrameWriter.cloudpickle==2.2.1 remains the only supported version.session.read.options where False Boolean values were incorrectly parsed as True in the generated file format.python-dateutil.Series.map when arg is a pandas Series or a
collections.abc.Mapping. No support for instances of dict that implement
__missing__ but are not instances of collections.defaultdict.DataFrame.align and Series.align for axis=1 and axis=None.pd.json_normalize.GroupBy.pct_change with axis=0, freq=None, and limit=None.DataFrameGroupBy.__iter__ and SeriesGroupBy.__iter__.np.sqrt, np.trunc, np.floor, numpy trig functions, np.exp, np.abs, np.positive and np.negative.DataFrame.__dataframe__().df.loc where setting a single column from a series results in unexpected None values.Added the following new functions in snowflake.snowpark.dataframe:
snowflake.snowpark.dataframe:
mapinclude_error to Session.query_history to record queries that have error during execution.Session.get_session_stage is used instead of raising SnowparkSQLException.Session.stored_procedure_profiler.set_active_profiler.DataFrame:
cache_resultIn expression were used in selects.AttributeError while calling Session.stored_procedure_profiler.get_output when Session.stored_procedure_profiler is disabled.protobuf>=5.28 and tzlocal at runtime.protoc-wheel-0 for the development profile.snowflake-connector-python>=3.12.0, <4.0.0 (was >=3.10.0).modin from 0.28.1 to 0.30.1.pandas 2.2.x versions.Index.to_numpy.DataFrame.align and Series.align for axis=0.size in GroupBy.aggregate, DataFrame.aggregate, and Series.aggregate.snowflake.snowpark.functions.windowpd.read_pickle (Uses native pandas for processing).pd.read_html (Uses native pandas for processing).pd.read_xml (Uses native pandas for processing)."size" and len in GroupBy.aggregate, DataFrame.aggregate, and Series.aggregate.Series.str.len.pd.DataFrame([0]).agg(np.mean)) would fail to transpose the result.DataFrame.dropna() would:
subset (e.g. []) as if it specified all columns instead of no columns.TypeError for a scalar subset instead of filtering on just that column.ValueError for a subset of type pandas.Index instead of filtering on the columns in the index.TableNotFoundError when using dynamic pivot in notebook environment.snowflake.snowpark.functions module.snowflake.snowpark.functions.any_valueTable.update could not handle VariantType, MapType, and ArrayType data types.DataFrame.join, causing errors when selecting columns from a joined DataFrame.Table.update and Table.merge could fail if the target table's index was not the default RangeIndex.Deprecated warnings will be triggered when using snowpark-python with Python 3.8. For more details, please refer to https://docs.snowflake.com/en/deve…
Session class to be thread-safe. This allows concurrent DataFrame transformations, DataFrame actions, UDF and stored procedure registration, and concurrent file uploads when using the same Session object.
FEATURE_THREAD_SAFE_PYTHON_SESSION to True for account.DataFrame.queries API are not deterministic, and may be different when DataFrame actions are executed. This does not affect explicit user-created temporary tables.session.lineage.trace API.copy_grants parameter when registering UDxF and stored procedures.DataFrameWriter to support daisy-chaining:
optionoptionspartition_bysnowflake_cortex_summarize.snowflake.snowpark.functions.array_remove it is now possible to use in python.df.sort().limit() and df.limit().sort() generates the same query with sort in front of limit. Now, df.limit().sort() will generate query that reads df.limit().sort().df.limit().sort(), because limit stops table scanning as soon as the number of records is satisfied.DataFrame.analytics.time_series_agg function to handle multiple data points in same sliding interval.np.subtract, np.multiply, np.divide, and np.true_divide.__array_ufunc__.np.float_power, np.mod, np.remainder, np.greater, np.greater_equal, np.less, np.less_equal, np.not_equal, and np.equal.np.log, np.log2, and np.log10DataFrameGroupBy.bfill, SeriesGroupBy.bfill, DataFrameGroupBy.ffill, and SeriesGroupBy.ffill.on parameter with Resampler.value_counts().snowflake_cortex_summarize.DataFrame.attrs and Series.attrs.DataFrame.style.head and iloc when the row key is a slice.tz_convert and tz_localize in Series, DataFrame, Series.dt, and DatetimeIndex.tz_convert and tz_localize in Series, DataFrame, Series.dt, and DatetimeIndex to specify the supported timezone formats.df.apply and series.apply ( as well as map and applymap ) when using snowpark functions. This allows for some position independent compatibility between apply and functions where the first argument is not a pandas object.iloc and iat when the row key is a scalar.iterrows.Series.map to reflect the unsupported features.np.may_share_memory which is used internally by many scikit-learn functions. This method will always return false when called with a Snowpark pandas object.DataFrame and Series pct_change() would raise TypeError when input contained timedelta columns.replace() would sometimes propagate Timedelta types incorrectly through replace(). Instead raise NotImplementedError for replace() on Timedelta.DataFrame and Series round() would raise AssertionError for Timedelta columns. Instead raise NotImplementedError for round() on Timedelta.reindex fails when the new index is a Series with non-overlapping types from the original index.__getitem__ on a DataFrameGroupBy object always returned a DataFrameGroupBy object if as_index=False.NotImplementedError.DataFrame.shift() on axis=0 and axis=1 would fail to propagate timedelta types.DataFrame.abs(), DataFrame.__neg__(), DataFrame.stack(), and DataFrame.unstack() now raise NotImplementedError for timedelta inputs instead of failing to propagate timedelta types.DataFrame.alias raises KeyError for input column name.to_csv on Snowflake stage fails when data contains empty strings.Added the following new functions in snowflake.snowpark.functions:
snowflake.snowpark.functions:
make_intervalWindow.range_between() when the order by column is TIMESTAMP or DATE type.thread_id to QueryRecord to track the thread id submitting the query history.Session.stored_procedure_profiler.'NoneType' has no len() when trying to read default values from function.TimedeltaIndex.mean method.Timedelta columns on axis=0 with agg or aggregate.by, left_by, right_by, left_index, and right_index for pd.merge_asof.include_describe to Session.query_history.DatetimeIndex.mean and DatetimeIndex.std methods.Resampler.asfreq, Resampler.indices, Resampler.nunique, and Resampler.quantile.resample frequency W, ME, YE with closed = "left".DataFrame.rolling.corr and Series.rolling.corr for pairwise = False and int window.window and min_periods = None for Rolling.DataFrameGroupBy.fillna and SeriesGroupBy.fillna.Series and DataFrame objects with the lazy Index object as data, index, and columns arguments.Series and DataFrame objects with index and column values not present in DataFrame/Series data.pd.read_sas (Uses native pandas for processing).rolling().count() and expanding().count() to Timedelta series and columns.tz in both pd.date_range and pd.bdate_range.Series.items.errors="ignore" in pd.to_datetime.DataFrame.tz_localize and Series.tz_localize.DataFrame.tz_convert and Series.tz_convert.sin) in Series.map, Series.apply, DataFrame.apply and DataFrame.applymap.to_pandas to persist the original timezone offset for TIMESTAMP_TZ type.dtype results for TIMESTAMP_TZ type to show correct timezone offset.dtype results for TIMESTAMP_LTZ type to show correct timezone.numeric_only for groupby aggregations.sort_values.convert_dtype in Series.apply.Index object created from a Series/DataFrame incorrectly updates the Series/DataFrame's index name after an inplace update has been applied to the original Series/DataFrame.SettingWithCopyWarning that sometimes appeared when printing Timedelta columns.inplace argument for Series objects derived from other Series objects.Series.sort_values failed if series name overlapped with index column name.Timedelta index levels to integer column levels.Resampler methods on timedelta columns would produce integer results.pd.to_numeric() would leave Timedelta inputs as Timedelta instead of converting them to integers.loc set when setting a single row, or multiple rows, of a DataFrame with a Series value.This is a re-release of 1.22.0. Please refer to the 1.22.0 release notes for detailed release content.
This is a re-release of 1.22.0. Please refer to the 1.22.0 release notes for detailed release content.
snowflake.snowpark.functions:
array_removelnSession.write_pandas by making use_logical_type option more explicit.DataFrameWriter.save_as_table:
enable_schema_evolutiondata_retention_timemax_data_extension_timechange_trackingcopy_grantsiceberg_config A dicitionary that can hold the following iceberg configuration options:
external_volumecatalogbase_locationcatalog_syncstorage_serialization_policyDataFrameWriter.copy_into_table:
iceberg_config A dicitionary that can hold the following iceberg configuration options:
external_volumecatalogbase_locationcatalog_syncstorage_serialization_policyDataFrame.create_or_replace_dynamic_table:
moderefresh_modeinitializeclustering_keysis_transientdata_retention_timemax_data_extension_timesession.read.csv that caused an error when setting PARSE_HEADER = True in an externally defined file format.session.get_session_stage that referenced a non-existing stage after switching database or schema.DataFrame.to_snowpark_pandas without explicitly initializing the Snowpark pandas plugin caused an error.explode function in dynamic table creation caused a SQL compilation error due to improper boolean type casting on the outer parameter.Index.identical.DataFrameWriter.save_as_table incorrectly handled DataFrames containing only a subset of columns from the existing table.to_timestamp does not set the default timezone of the column datatype.Timedelta type, including the following features. Snowpark pandas will raise NotImplementedError for unsupported Timedelta use cases.
copy, cache_result, shift, sort_index, assign, bfill, ffill, fillna, compare, diff, drop, dropna, duplicated, empty, equals, insert, isin, isna, items, iterrows, join, len, mask, melt, merge, nlargest, nsmallest, to_pandas.astype.NotImplementedError will be raised for the rest of methods that do not support Timedelta.Timedelta.Timedelta values.Timedelta values and numeric values.TimedeltaIndex.pd.to_timedelta.GroupBy aggregations min, max, mean, idxmax, idxmin, std, sum, median, count, any, all, size, nunique, head, tail, aggregate.GroupBy filtrations first and last.TimedeltaIndex attributes: days, seconds, microseconds and nanoseconds.diff with timestamp columns on axis=0 and axis=1TimedeltaIndex methods: ceil, floor and round.TimedeltaIndex.total_seconds method.Series.dt.round.DatetimeIndex.Index.name, Index.names, Index.rename, and Index.set_names.Index.__repr__.DatetimeIndex.month_name and DatetimeIndex.day_name.Series.dt.weekday, Series.dt.time, and DatetimeIndex.time.Index.min and Index.max.pd.merge_asof.Series.dt.normalize and DatetimeIndex.normalize.Index.is_boolean, Index.is_integer, Index.is_floating, Index.is_numeric, and Index.is_object.DatetimeIndex.round, DatetimeIndex.floor and DatetimeIndex.ceil.Series.dt.days_in_month and Series.dt.daysinmonth.DataFrameGroupBy.value_counts and SeriesGroupBy.value_counts.Series.is_monotonic_increasing and Series.is_monotonic_decreasing.Index.is_monotonic_increasing and Index.is_monotonic_decreasing.pd.crosstab.pd.bdate_range and included business frequency support (B, BME, BMS, BQE, BQS, BYE, BYS) for both pd.date_range and pd.bdate_range.Index objects as labels in DataFrame.reindex and Series.reindex.Series.dt.days, Series.dt.seconds, Series.dt.microseconds, and Series.dt.nanoseconds.DatetimeIndex from an Index of numeric or string type.Timedelta objects.Series.dt.total_seconds method.quoted_identifier_to_snowflake_type to avoid making metadata queries if the types have been cached locally.pd.to_datetime to handle all local input cases.NotImplementedError for Index bitwise operators.Index.names is set to a non-like-like object.pd.read_snowflake include the creation reason when temp table creation is triggered.DataFrame.set_index, or setting DataFrame.index or Series.index by avoiding checks require eager evaluation. As a consequence, when the new index that does not match the current Series/DataFrame object length, a ValueError is no longer raised. Instead, when the Series/DataFrame object is longer than the provided index, the Series/DataFrame's new index is filled with NaN values for the "extra" elements. Otherwise, the extra values in the provided index are ignored.pd.Timedelta scalars.Series.dt.isocalendar using a named Seriesinplace argument for Series objects derived from DataFrame columns.Series.reindex and DataFrame.reindex did not update the result index's name correctly.Series.take did not error when axis=1 was specified.Added the following new functions in snowflake.snowpark.functions:
snowflake.snowpark.functions:
array_removelnSession.write_pandas by making use_logical_type option more explicit.DataFrameWriter.save_as_table:
enable_schema_evolutiondata_retention_timemax_data_extension_timechange_trackingcopy_grantsiceberg_config A dicitionary that can hold the following iceberg configuration options:
external_volumecatalogbase_locationcatalog_syncstorage_serialization_policyDataFrameWriter.copy_into_table:
iceberg_config A dicitionary that can hold the following iceberg configuration options:
external_volumecatalogbase_locationcatalog_syncstorage_serialization_policyDataFrame.create_or_replace_dynamic_table:
moderefresh_modeinitializeclustering_keysis_transientdata_retention_timemax_data_extension_timesession.read.csv that caused an error when setting PARSE_HEADER = True in an externally defined file format.session.get_session_stage that referenced a non-existing stage after switching database or schema.DataFrame.to_snowpark_pandas without explicitly initializing the Snowpark pandas plugin caused an error.explode function in dynamic table creation caused a SQL compilation error due to improper boolean type casting on the outer parameter.Index.identical.DataFrameWriter.save_as_table incorrectly handled DataFrames containing only a subset of columns from the existing table.to_timestamp does not set the default timezone of the column datatype.Timedelta type, including the following features. Snowpark pandas will raise NotImplementedError for unsupported Timedelta use cases.
copy, cache_result, shift, sort_index, assign, bfill, ffill, fillna, compare, diff, drop, dropna, duplicated, empty, equals, insert, isin, isna, items, iterrows, join, len, mask, melt, merge, nlargest, nsmallest, to_pandas.astype.NotImplementedError will be raised for the rest of methods that do not support Timedelta.Timedelta.Timedelta values.Timedelta values and numeric values.TimedeltaIndex.pd.to_timedelta.GroupBy aggregations min, max, mean, idxmax, idxmin, std, sum, median, count, any, all, size, nunique, head, tail, aggregate.GroupBy filtrations first and last.TimedeltaIndex attributes: days, seconds, microseconds and nanoseconds.diff with timestamp columns on axis=0 and axis=1TimedeltaIndex methods: ceil, floor and round.TimedeltaIndex.total_seconds method.Series.dt.round.DatetimeIndex.Index.name, Index.names, Index.rename, and Index.set_names.Index.__repr__.DatetimeIndex.month_name and DatetimeIndex.day_name.Series.dt.weekday, Series.dt.time, and DatetimeIndex.time.Index.min and Index.max.pd.merge_asof.Series.dt.normalize and DatetimeIndex.normalize.Index.is_boolean, Index.is_integer, Index.is_floating, Index.is_numeric, and Index.is_object.DatetimeIndex.round, DatetimeIndex.floor and DatetimeIndex.ceil.Series.dt.days_in_month and Series.dt.daysinmonth.DataFrameGroupBy.value_counts and SeriesGroupBy.value_counts.Series.is_monotonic_increasing and Series.is_monotonic_decreasing.Index.is_monotonic_increasing and Index.is_monotonic_decreasing.pd.crosstab.pd.bdate_range and included business frequency support (B, BME, BMS, BQE, BQS, BYE, BYS) for both pd.date_range and pd.bdate_range.Index objects as labels in DataFrame.reindex and Series.reindex.Series.dt.days, Series.dt.seconds, Series.dt.microseconds, and Series.dt.nanoseconds.DatetimeIndex from an Index of numeric or string type.Timedelta objects.Series.dt.total_seconds method.DataFrame.apply(axis=0).Series.dt.tz_convert and Series.dt.tz_localize.DatetimeIndex.tz_convert and DatetimeIndex.tz_localize.quoted_identifier_to_snowflake_type to avoid making metadata queries if the types have been cached locally.pd.to_datetime to handle all local input cases.NotImplementedError for Index bitwise operators.Index.names is set to a non-like-like object.pd.read_snowflake include the creation reason when temp table creation is triggered.DataFrame.set_index, or setting DataFrame.index or Series.index by avoiding checks require eager evaluation. As a consequence, when the new index that does not match the current Series/DataFrame object length, a ValueError is no longer raised. Instead, when the Series/DataFrame object is longer than the provided index, the Series/DataFrame's new index is filled with NaN values for the "extra" elements. Otherwise, the extra values in the provided index are ignored.NotImplementedError when ambiguous/nonexistent are non-string in ceil/floor/round.pd.Timedelta scalars.Series.dt.isocalendar using a named Seriesinplace argument for Series objects derived from DataFrame columns.Series.reindex and DataFrame.reindex did not update the result index's name correctly.Series.take did not error when axis=1 was specified.Fixed a bug where using to_pandas_batches with async jobs caused an error due to improper handling of waiting for asynchronous query completion.
to_pandas_batches with async jobs caused an error due to improper handling of waiting for asynchronous query completion.Added support for snowflake.snowpark.testing.assert_dataframe_equal that is a utility function to check the equality of two Snowpark DataFrames.
snowflake.snowpark.testing.assert_dataframe_equal that is a utility function to check the equality of two Snowpark DataFrames.INFER_SCHEMA options to DataFrameReader via INFER_SCHEMA_OPTIONS.parameters parameter to Column.rlike and Column.regexp.df.cache_result() in the current session, when the DataFrame is no longer referenced (i.e., gets garbage collected). It is still an experimental feature not enabled by default, and can be enabled by setting session.auto_clean_up_temp_table_enabled to True.fmt parameter of snowflake.snowpark.functions.to_date.* column has an incorrect subquery.DataFrame.to_pandas_batches where the iterator could throw an error if certain transformation is made to the pandas dataframe due to wrong isolation level.DataFrame.lineage.trace to split the quoted feature view's name and version correctly.Column.isin that caused invalid sql generation when passed an empty list.rankdense_rankpercent_rankcume_distntiledatediffarray_aggrlike and regexp changes above.ignore_nulls properly.DataFrame.backfill, DataFrame.bfill, Series.backfill, and Series.bfill.DataFrame.compare and Series.compare with default parameters.Series.dt.microsecond and Series.dt.nanosecond.Index.is_unique and Index.has_duplicates.Index.equals.Index.value_counts.Series.dt.day_name and Series.dt.month_name.df.index[:10].DataFrame.unstack and Series.unstack.DataFrame.asfreq and Series.asfreq.Series.dt.is_month_start and Series.dt.is_month_end.Index.all and Index.any.Series.dt.is_year_start and Series.dt.is_year_end.Series.dt.is_quarter_start and Series.dt.is_quarter_end.DatetimeIndex.Series.argmax and Series.argmin.Series.dt.is_leap_year.DataFrame.items.Series.dt.floor and Series.dt.ceil.Index.reindex.DatetimeIndex properties: year, month, day, hour, minute, second, microsecond,
nanosecond, date, dayofyear, day_of_year, dayofweek, day_of_week, weekday, quarter,
is_month_start, is_month_end, is_quarter_start, is_quarter_end, is_year_start, is_year_end
and is_leap_year.Resampler.fillna and Resampler.bfill.Timedelta type, including creating Timedelta columns and to_pandas.Index.argmax and Index.argmin.SnowflakeQueryCompiler.is_series_like method.Dataframe.columns now returns native pandas Index object instead of Snowpark Index object.query_compiler argument in Index constructor to create Index from query compiler.pd.to_datetime now returns a DatetimeIndex object instead of a Series object.pd.date_range now returns a DatetimeIndex object instead of a Series object.pivot_table raise NotImplementedError instead of KeyError.Series.drop_duplicates and DataFrame.drop_duplicates when called after sort_values.Index.to_frame where the result frame's column name may be wrong where name is unspecified.Series.reset_index(drop=True) where the result name may be wrong.Groupby.first/last ordering by the correct columns in the underlying window expression.Added distributed tracing using open telemetry APIs for table stored procedure function in DataFrame:
DataFrame:
_execute_and_get_query_idarrays_zip function.df._in by avoiding unnecessary cast for numeric values. You can enable this optimization by setting session.eliminate_numeric_sql_value_cast_enabled = True.write_pandas when the target table does not exist and auto_create_table=False.format_json to the Session.SessionBuilder.app_name function that sets the app name in the Session.query_tag in JSON format. By default, this parameter is set to False.lag(x, 0) was incorrect and failed with error message argument 1 to function LAG needs to be constant, found 'SYSTEM$NULL_TO_FIXED(null)'.patch function when registering a mocked function:
distinct allows an alternate function to be specified for when a sql function should be distinct.pass_column_index passes a named parameter column_index to the mocked function that contains the pandas.Index for the input data.pass_row_index passes a named parameter row_index to the mocked function that is the 0 indexed row number the function is currently operating on.pass_input_data passes a named parameter input_data to the mocked function that contains the entire input dataframe for the current expression.column_order parameter to method DataFrameWriter.save_as_table.DataFrameGroupBy.all, SeriesGroupBy.all, DataFrameGroupBy.any, and SeriesGroupBy.any.DataFrame.nlargest, DataFrame.nsmallest, Series.nlargest and Series.nsmallest.replace and frac > 1 in DataFrame.sample and Series.sample.read_excel (Uses local pandas for processing)Series.at, Series.iat, DataFrame.at, and DataFrame.iat.Series.dt.isocalendar.Series.case_when except when condition or replacement is callable.Index and its APIs.DataFrame.assign.DataFrame.stack.DataFrame.pivot and pd.pivot.DataFrame.to_csv and Series.to_csv.Series.str.translate where the values in the table are single-codepoint strings.DataFrame.corr.df.plot() and series.plot() to be called, materializing the data into the local clientDataFrameGroupBy and SeriesGroupBy aggregations first and lastDataFrameGroupBy.get_group.limit parameter when method parameter is used in fillna.Series.str.translate where the values in the table are single-codepoint strings.DataFrame.corr.DataFrame.equals and Series.equals.DataFrame.reindex and Series.reindex.Index.astype.Index.unique and Index.nunique.DataFrame or Series with dtype=np.uint64.values is set to index when index and columns contain all columns in DataFrame during pivot_table.Index.copy()dtype, values, item(), tolist(), to_series() and to_frame()pd.pivot_table and DataFrame.pivot_table.inplace parameter in DataFrame.sort_index and Series.sort_index.Added support for to_boolean function.
to_boolean function.Index and its APIs.RecursionError: maximum recursion depth exceeded when the DataFrame has more than 500 columns.AsyncJob.result("no_result") doesn't wait for the query to finish execution.strict parameter when registering UDFs and Stored Procedures.DateType raises AttributeError.to_char that raises IndexError when incoming column has nonconsecutive row index.CaseExpr expressions that raises IndexError when incoming column has nonconsecutive row index.Column.like that raises IndexError when incoming column has nonconsecutive row index.iff.DataFrame.pct_change and Series.pct_change without the freq and limit parameters.Series.str.get.Series.dt.dayofweek, Series.dt.day_of_week, Series.dt.dayofyear, and Series.dt.day_of_year.Series.str.__getitem__ (Series.str[...]).Series.str.lstrip and Series.str.rstrip.DataFrameGroupby.size and SeriesGroupby.size.DataFrame.expanding and Series.expanding for aggregations count, sum, min, max, mean, std, and var with axis=0.DataFrame.rolling and Series.rolling for aggregation count with axis=0.Series.str.match.DataFrame.resample and Series.resample for aggregation size.DataFrame.describe on a frame with duplicate columns of differing dtypes could cause an error or incorrect results.DataFrame.rolling and Series.rolling so window=0 now throws NotImplementedError instead of ValueErrorDataFrame.aggregate and Series.aggregate with axis=0.pd.read_csv reads using the native pandas CSV parser, then uploads data to snowflake using parquet. This enables most of the parameters supported by read_csv including date parsing and numeric conversions. Uploading via parquet is roughly twice as fast as uploading via CSV.pd.Index directly in Snowpark pandas. Support for pd.Index as a first-class component of Snowpark pandas is coming soon.len, shape, size, empty, to_pandas() and names. For df.index, Snowpark pandas creates a lazy index object.df.columns, Snowpark pandas supports a non-lazy version of an Index since the data is already stored locally.to_boolean function.RecursionError: maximum recursion depth exceeded when the DataFrame has more than 500 columns.AsyncJob.result("no_result") doesn't wait for the query to finish execution.strict parameter when registering UDFs and Stored Procedures.DateType raises AttributeError.to_char that raises IndexError when incoming column has nonconsecutive row index.CaseExpr expressions that raises IndexError when incoming column has nonconsecutive row index.Column.like that raises IndexError when incoming column has nonconsecutive row index.iff.DataFrame.pct_change and Series.pct_change without the freq and limit parameters.Series.str.get.Series.dt.dayofweek, Series.dt.day_of_week, Series.dt.dayofyear, and Series.dt.day_of_year.Series.str.__getitem__ (Series.str[...]).Series.str.lstrip and Series.str.rstrip.DataFrameGroupBy.size and SeriesGroupBy.size.DataFrame.expanding and Series.expanding for aggregations count, sum, min, max, mean, std, var, and sem with axis=0.DataFrame.rolling and Series.rolling for aggregation count with axis=0.Series.str.match.DataFrame.resample and Series.resample for aggregations size, first, and last.DataFrameGroupBy.all, SeriesGroupBy.all, DataFrameGroupBy.any, and SeriesGroupBy.any.DataFrame.nlargest, DataFrame.nsmallest, Series.nlargest and Series.nsmallest.replace and frac > 1 in DataFrame.sample and Series.sample.read_excel (Uses local pandas for processing)Series.at, Series.iat, DataFrame.at, and DataFrame.iat.Series.dt.isocalendar.Series.case_when except when condition or replacement is callable.Index and its APIs.DataFrame.assign.DataFrame.stack.DataFrame.pivot and pd.pivot.DataFrame.to_csv and Series.to_csv.Index.T.DataFrame.describe on a frame with duplicate columns of differing dtypes could cause an error or incorrect results.DataFrame.rolling and Series.rolling so window=0 now throws NotImplementedError instead of ValueErrorDataFrame.aggregate and Series.aggregate with axis=0.pd.read_csv reads using the native pandas CSV parser, then uploads data to snowflake using parquet. This enables most of the parameters supported by read_csv including date parsing and numeric conversions. Uploading via parquet is roughly twice as fast as uploading via CSV.pd.Index directly in Snowpark pandas. Support for pd.Index as a first-class component of Snowpark pandas is coming soon.len, shape, size, empty, to_pandas() and names. For df.index, Snowpark pandas creates a lazy index object.df.columns, Snowpark pandas supports a non-lazy version of an Index since the data is already stored locally.Added DataFrame.cache_result and Series.cache_result methods for users to persist DataFrames and Series to a temporary table lasting the duration of t
DataFrame.cache_result and Series.cache_result methods for users to persist DataFrames and Series to a temporary table lasting the duration of the session to improve latency of subsequent operations.DataFrame.pivot_table with no index parameter, as well as for margins parameter.DataFrame.shift/Series.shift/DataFrameGroupBy.shift/SeriesGroupBy.shift to match pandas 2.2.1. Snowpark pandas does not yet support the newly-added suffix argument, or sequence values of periods.Series.str.split.Series.str.*).csv and json:
FalseUTF8DataFrame.analytics.moving_agg and DataFrame.analytics.cumulative_agg_agg.if_not_exists parameter during UDF and stored procedure registration.* to fail.date_add was unable to handle some numeric types.TimestampType casting resulted in incorrect data.DecimalType data to have incorrect precision in some cases.IndexError.to_timestamp_ntz can not handle None data.DataFrame.with_column_renamed ignores attributes from parent DataFrames after join operations.Column.equal_nan where null data is handled incorrectly.DataFrame.drop ignore attributes from parent DataFrames after join operations.date_part where Column type is set wrong.DataFrameWriter.save_as_table does not raise exceptions when inserting null data into non-nullable columns.DataFrameWriter.save_as_table where
pyarrow as it is not used.Column.cast, adding support for casting to boolean and all integral types.is_permanent and anonymous options in UDFs and stored procedures registration to make it more clear that those features are not yet supported.NotImplementedError instead of warnings and unclear error information.Added support to add a comment on tables and views using the functions listed below:
DataFrameWriter.save_as_tableDataFrame.create_or_replace_viewDataFrame.create_or_replace_temp_viewDataFrame.create_or_replace_dynamic_table{"infer_schema": True} when reading CSV file without specifying its schema.to_timestamp_ltz, to_timestamp_ntz, to_timestamp_tz and to_timestamp.to_char.snowflake.snowpark.mock.exceptions.SnowparkLocalTestingException.sys.path during the clean-up step.Session.get_current_[schema|database|role|user|account|warehouse] returns upper-cased identifiers when identifiers are quoted.substr and substring can not handle 0-based start_expr.SnowparkLocalTestingException in error cases which is on par with SnowparkSQLException raised in non-local execution.Session.write_pandas method that NotImplementError will be raised when called.Added snowflake.snowpark.Session.lineage.trace to explore data lineage of Snowflake objects.
to_date.None value in an arithmetic calculation, the output should remain None instead of math.nan.sum and covar_pop that when there is math.nan in the data, the output should also be math.nan.DataFrame.to_pandas should take Snowflake numeric types with precision 38 as int64.Added truncate save mode in DataFrameWrite to overwrite existing tables by truncating the underlying table instead of dropping it.
truncate save mode in DataFrameWrite to overwrite existing tables by truncating the underlying table instead of dropping it.DataFrame into one or more files in a stage:
DataFrame.write.jsonDataFrame.write.csvDataFrame.write.parquetDataFrame and DataFrameWriter:
snowflake.snowpark.Session.file.get and snowflake.snowpark.Session.file.get_streamcomment.session.cte_optimization_enabled to True.statement_params was not passed to query executions that register stored procedures and user defined functions.snowflake.snowpark.Session.file.get_stream to fail for quoted stage locations.utils.py might raise AttributeError in case the underlying module can not be found.to_time.Session.builder.getOrCreate should return the created mock session.Added support for creating vectorized UDTFs with process method.
process method.SnowflakePlanBuilder that save_as_table does not filter column that name start with '$' and follow by number correctly.field_optionally_enclosed_by is specified.pattern is a Column.KeyError when updating null values in the rows.DataFrame.collect.count_distinct does not work correctly when counting.TypeError.DataFrameReader to raise FileNotFound error when reading a path that does not exist or when there are no files under the path.Added support for an optional date_part argument in function last_day
date_part argument in function last_daySessionBuilder.app_name will set the query_tag after the session is created.DataFrame.to_local_iterator where the iterator could yield wrong results if another query is executed before the iterator finishes due to wrong isolation level. For details, please see #945.Session.range returns empty result when the range is large.date_part argument in function last_day.SessionBuilder.app_name will set the query_tag after the session is created.DataFrame.to_local_iterator where the iterator could yield wrong results if another query is executed before the iterator finishes due to wrong isolation level. For details, please see #945.Session.range returns empty result when the range is large.Use split_blocks=True by default during to_pandas conversion, for optimal memory allocation. This parameter is passed to pyarrow.Table.to_pandas, whic
split_blocks=True by default during to_pandas conversion, for optimal memory allocation. This parameter is passed to pyarrow.Table.to_pandas, which enables PyArrow to split the memory allocation into smaller, more manageable blocks instead of allocating a single contiguous block. This results in better memory management when dealing with larger datasets.DataFrame.to_pandas that caused an error when evaluating on a Dataframe with an IntergerType column with null values.Exposed statement_params in StoredProcedure.__call__.
statement_params in StoredProcedure.__call__.Session.add_import.
chunk_size: The number of bytes to hash per chunk of the uploaded files.whole_file_hash: By default only the first chunk of the uploaded import is hashed to save time. When this is set to True each uploaded file is fully hashed instead.external_access_integrations and secrets when creating a UDAF from Snowpark Python to allow integration with external access.Session.append_query_tag. Allows an additional tag to be added to the current query tag by appending it as a comma separated value.Session.update_query_tag. Allows updates to a JSON encoded dictionary query tag.SessionBuilder.getOrCreate will now attempt to replace the singleton it returns when token expiration has been detected.snowflake.snowpark.functions:
array_exceptcreate_mapsign/signumDataFrame.analytics:
moving_agg function in DataFrame.analytics to enable moving aggregations like sums and averages with multiple window sizes.cummulative_agg function in DataFrame.analytics to enable moving aggregations like sums and averages with multiple window sizes.Fixed a bug in DataFrame.na.fill that caused Boolean values to erroneously override integer values.
Fixed a bug in Session.create_dataframe where the Snowpark DataFrames created using pandas DataFrames were not inferring the type for timestamp columns correctly. The behavior is as follows:
LongType(), but will now be correctly maintained as timestamp values and be inferred as TimestampType(TimestampTimeZone.NTZ).TimestampType(TimestampTimeZone.NTZ) and loose timezone information but will now be correctly inferred as TimestampType(TimestampTimeZone.LTZ) and timezone information is retained correctly.PYTHON_SNOWPARK_USE_LOGICAL_TYPE_FOR_CREATE_DATAFRAME to revert back to old behavior. It is recommended that you update your code to align with correct behavior because the parameter will be removed in the future.Fixed a bug that DataFrame.to_pandas gets decimal type when scale is not 0, and creates an object dtype in pandas. Instead, we cast the value to a float64 type.
Fixed bugs that wrongly flattened the generated SQL when one of the following happens:
DataFrame.filter() is called after DataFrame.sort().limit().DataFrame.sort() or filter() is called on a DataFrame that already has a window function or sequence-dependent data generator column.
For instance, df.select("a", seq1().alias("b")).select("a", "b").sort("a") won't flatten the sort clause anymore.DataFrame.limit(). For instance, df.limit(10).select(row_number().over()) won't flatten the limit and select in the generated SQL.Fixed a bug where aliasing a DataFrame column raised an error when the DataFame was copied from another DataFrame with an aliased column. For instance,
df = df.select(col("a").alias("b"))
df = copy(df)
df.select(col("b").alias("c")) # threw an error. Now it's fixed.
Fixed a bug in Session.create_dataframe that the non-nullable field in a schema is not respected for boolean type. Note that this fix is only effective when the user has the privilege to create a temp table.
Fixed a bug in SQL simplifier where non-select statements in session.sql dropped a SQL query when used with limit().
Fixed a bug that raised an exception when session parameter ERROR_ON_NONDETERMINISTIC_UPDATE is true.
to_pandas operation, we rely on GS precision value to fix precision issues for large integer values. This may affect users where a column that was earlier returned as int8 gets returned as int64. Users can fix this by explicitly specifying precision values for their return column.Session.call in case of table stored procedures where running Session.call would not trigger stored procedure unless a collect() operation was performed.StoredProcedureRegistration will now automatically add snowflake-snowpark-python as a package dependency. The added dependency will be on the client's local version of the library and an error is thrown if the server cannot support that version.Fixed a bug that numpy should not be imported at the top level of mock module.
snowflake.snowpark.functions:
from_utc_timestampto_utc_timestampAdd the conn_error attribute to SnowflakeSQLException that stores the whole underlying exception from snowflake-connector-python.
Add the conn_error attribute to SnowflakeSQLException that stores the whole underlying exception from snowflake-connector-python.
Added support for RelationalGroupedDataframe.pivot() to access pivot in the following pattern Dataframe.group_by(...).pivot(...).
Added experimental feature: Local Testing Mode, which allows you to create and operate on Snowpark Python DataFrames locally without connecting to a Snowflake account. You can use the local testing framework to test your DataFrame operations locally, on your development machine or in a CI (continuous integration) pipeline, before deploying code changes to your account.
Added support for arrays_to_object new functions in snowflake.snowpark.functions.
Added support for the vector data type.
cloudpickle==2.2.1snowflake-connector-python to 3.4.0.session.read.with_metadata creates inconsistent table when doing df.write.save_as_table.Added support for managing case sensitivity in DataFrame.to_local_iterator().
DataFrame.to_local_iterator().input_names in UDTFRegistration.register/register_file and functions.pandas_udtf. By default, RelationalGroupedDataFrame.applyInPandas will infer the column names from current dataframe schema.sql_error_code and raw_message attributes to SnowflakeSQLException when it is caused by a SQL exception.DataFrame.to_pandas() where converting snowpark dataframes to pandas dataframes was losing precision on integers with more than 19 digits.session.add_packages can not handle requirement specifier that contains project name with underscore and version.DataFrame.limit() when offset is used and the parent DataFrame uses limit. Now the offset won't impact the parent DataFrame's limit.DataFrame.write.save_as_table where dataframes created from read api could not save data into snowflake because of invalid column name $1.date_format:
format argument changed from optional to required.normal, zipf, uniform, seq1, seq2, seq4, seq8) function is used, the sort and filter operation will no longer be flattened when generating the query.Added support for the Python 3.11 runtime environment.
typing-extensions.Dataframe.writer.save_as_table which does not need insert permission for writing tables.PythonObjJSONEncoder json-serializable objects for ARRAY and OBJECT literals.Added support for VOLATILE/IMMUTABLE keyword when registering UDFs.
DataFrame.save_as_table.Iterable objects input for schema when creating dataframes using Session.create_dataframe.DataFrame.session to return a Session object.Session.session_id to return an integer that represents session ID.Session.connection to return a SnowflakeConnection object .snowflake-connector-python to 3.2.0.ValueError even when compatible package version were added in session.add_packages.register_from_file.invalid_identifier error.DataFrame.copy disables SQL simplfier for the returned copy.session.sql().select() would fail if any parameters are specified to session.sql().Added parameters external_access_integrations and secrets when creating a UDF, UDTF or Stored Procedure from Snowpark Python to allow integration with
external_access_integrations and secrets when creating a UDF, UDTF or Stored Procedure from Snowpark Python to allow integration with external access.snowflake.snowpark.functions:
array_flattenflattenapply_in_pandas in snowflake.snowpark.relational_grouped_dataframe.Session.replicate_local_environment.session.create_dataframe fails to properly set nullable columns where nullability was affected by order or data was given.DataFrame.select could not identify and alias columns in presence of table functions when output columns of table function overlapped with columns in dataframe.is_permanent=False will now create temporary objects even when stage_name is provided. The default value of is_permanent is False which is why if this value is not explicitly set to True for permanent objects, users will notice a change in behavior.types.StructField now enquotes column identifier by default.Your coding agent can read these notes before it upgrades. Set up the MCP server →