NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1517 most downloaded on PyPI
A light-weight and flexible data validation and testing tool for statistical data objects.
Last release 17 days ago
01 Sep 2026
Ships fairly regularly
a new release about every 4 weeks
Nearly every release is documented
notes for 58 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
8 years old
125 releases · first in 2018
Updates the maximum python version in setup.py
Updates the maximum python version in setup.py
Support pandas 2 by @cosmicBboy in https://github.com/unionai-oss/pandera/pull/1175
default column value param by @kykyi in https://github.com/unionai-oss/pandera/pull/1136| ⚠️ Note |
|---|
Support for pandas 2 with dask, modin, and pyspark is currently not tested and is not guaranteed to work |
Shoutout to the new contributors! 📣
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.5...v0.15.0
One column per quarter.
Nothing published for this version
refactor: move error reasons to an enum by @DillonSteyl in https://github.com/unionai-oss/pandera/pull/1035
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.5...v0.14.5
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.4...v0.14.5
fix regression: handle schema.coerce_dtype correctly by @cosmicBboy in https://github.com/unionai-oss/pandera/pull/1119
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.3...v0.14.4
fix regression: all missing columns should be reported by @cosmicBboy in https://github.com/unionai-oss/pandera/pull/1117
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.2...v0.14.3
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.1...v0.14.2
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.1...v0.14.2
multimethod is a runtime dependency by @timkpaine in https://github.com/unionai-oss/pandera/pull/1112
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.14.0...v0.14.1
The main highlight of this release is that phase 1 of the Pandera internals re-write is complete 🎉🚀! This is a backwards-compatible re-write (unit tes
The main highlight of this release is that phase 1 of the Pandera internals re-write is complete 🎉🚀! This is a backwards-compatible re-write (unit tests FTW 😅) that should just work with your existing pandera code. Please submit bug reports if you encounter any regressions that weren't covered by the current test suite.
These PRs #913 #1109, and #1110 address https://github.com/unionai-oss/pandera/issues/381, and essentially decouples pandas-specific logic from the pandera schema specification. In summary:
pandera.api, containing:
pandera.api.basepandera.api.pandaspandera.api.checks.Check and pandera.api.hypotheses.Hypothesispandera.api.extensions to be able to register builtin and custom checks/hypothesespandera.backends, containing:
pandera.backends.basepandera.backends.pandasNow, all pandas-specific logic is isolated to specific modules, where support for additional non-pandas-compliant schema specifications and their associated backends can be implemented either as 1st-party-maintained libraries (see issues for supporting polars and ibis) or 3rd party libraries.
The bulk of the re-write is complete, however there are still some outstanding items:
pandera-ibis package, which will create a schema specification and validation backends for ibis data structures (see issue https://github.com/unionai-oss/pandera/issues/1105)coerce==True when for PydandticModels by @the-matt-morris in https://github.com/unionai-oss/pandera/pull/1011check_types: Bugfix/977 by @kr-hansen in https://github.com/unionai-oss/pandera/pull/995Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.13.4...v0.14.0
fix tz deprecation by @cristianmatache in https://github.com/unionai-oss/pandera/pull/972
Shoutout to all the contributors of this release 🎉
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.13.3...v0.13.4
Nothing published for this version
Fix decimal by @cosmicBboy in https://github.com/unionai-oss/pandera/pull/956
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.13.2...v0.13.3
fix modin tests by @cosmicBboy in https://github.com/unionai-oss/pandera/pull/955
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.13.1...v0.13.2
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.13.0...v0.13.1
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.13.0...v0.13.1
try pandera: add jupyterlite notebooks, add support for py3.7 (#951) @cosmicBboy
ordered=True in yaml schemas (#943) @dstumPYShout out to all the first-time contributors!
Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.12.0...v0.13.0
Nothing published for this version
Nothing published for this version
Support for Logical Data Types #798: These data types check the actual values of the data container at runtime to support data types like "URL", "Name
This release features:
"URL", "Name", etc.Check.unique_values_eq #858: Make sure that all of the values in the data container cover the entire domain of the specified finite set of values.ExtensionDtype path should follow documentation by @pepelovesvim in https://github.com/unionai-oss/pandera/pull/915Full Changelog: https://github.com/unionai-oss/pandera/compare/v0.11.0...v0.12.0
Just a little something for folks who prefer dark mode!
Just a little something for folks who prefer dark mode!
<img width="1328" alt="image" src="https://user-images.githubusercontent.com/2816689/166127024-5e473c9d-9d31-4775-b9cb-ccdab1edb4c4.png">
Nothing published for this version
Nothing published for this version
Nothing published for this version
The pandera [koalas](https://koalas.readthedocs.io/en/latest/index.html) integration has now been deprecated
pandera now supports pyspark dataframe validation via pyspark.pandasThe pandera koalas integration has now been deprecated
You can now pip install pandera[pyspark] and validate pyspark.pandas dataframes:
import pyspark.pandas as ps
import pandas as pd
import pandera as pa
from pandera.typing.pyspark import DataFrame, Series
class Schema(pa.SchemaModel):
state: Series[str]
city: Series[str]
price: Series[int] = pa.Field(in_range={"min_value": 5, "max_value": 20})
# create a pyspark.pandas dataframe that's validated on object initialization
df = DataFrame[Schema](
{
'state': ['FL','FL','FL','CA','CA','CA'],
'city': [
'Orlando',
'Miami',
'Tampa',
'San Francisco',
'Los Angeles',
'San Diego',
],
'price': [8, 12, 10, 16, 20, 18],
}
)
print(df)
PydanticModel DataType Enables Row-wise Validation with a pydantic modelPandera now supports row-wise validation by applying a pydantic model as a dataframe-level dtype:
from pydantic import BaseModel
import pandera as pa
class Record(BaseModel):
name: str
xcoord: str
ycoord: int
import pandas as pd
from pandera.engines.pandas_engine import PydanticModel
class PydanticSchema(pa.SchemaModel):
"""Pandera schema using the pydantic model."""
class Config:
"""Config with dataframe-level data type."""
dtype = PydanticModel(Record)
coerce = True # this is required, otherwise a SchemaInitError is raised
⚠️ Warning: This may lead to performance issues for very large dataframes.
Before this release there were only two conda packages: one to install pandera-core and another to install pandera (which would install all extras functionality)
The conda packaging now supports finer-grained control:
conda install -c conda-forge pandera-hypotheses # hypothesis checks
conda install -c conda-forge pandera-io # yaml/script schema io utilities
conda install -c conda-forge pandera-strategies # data synthesis strategies
conda install -c conda-forge pandera-mypy # enable static type-linting of pandas
conda install -c conda-forge pandera-fastapi # fastapi integration
conda install -c conda-forge pandera-dask # validate dask dataframes
conda install -c conda-forge pandera-pyspark # validate pyspark dataframes
conda install -c conda-forge pandera-modin # validate modin dataframes
conda install -c conda-forge pandera-modin-ray # validate modin dataframes with ray
conda install -c conda-forge pandera-modin-dask # validate modin dataframes with dask
pandera now integrates with fastapi. You can decorate app endpoint arguments with DataFrame[Schema] types and the endpoint will validate incoming and
pandera now integrates with fastapi. You can decorate app endpoint arguments with DataFrame[Schema] types and the endpoint will validate incoming and outgoing data.
from typing import Optional
from pydantic import BaseModel, Field
import pandera as pa
# schema definitions
class Transactions(pa.SchemaModel):
id: pa.typing.Series[int]
cost: pa.typing.Series[float] = pa.Field(ge=0, le=1000)
class Config:
coerce = True
class TransactionsOut(Transactions):
id: pa.typing.Series[int]
cost: pa.typing.Series[float]
name: pa.typing.Series[str]
class TransactionsDictOut(TransactionsOut):
class Config:
to_format = "dict"
to_format_kwargs = {"orient": "records"}
App endpoint example:
from fastapi import FastAPI, File
app = FastAPI()
@app.post("/transactions/", response_model=DataFrame[TransactionsDictOut])
def create_transactions(transactions: DataFrame[Transactions]):
output = transactions.assign(name="foo")
... # do other stuff, e.g. update backend database with transactions
return output
The class-based API now supports automatically deserializing/serializing pandas dataframes in the context of @pa.check_types-decorated functions, @pydantic.validate_arguments-decorated functions, and fastapi endpoint functions.
import pandera as pa
from pandera.typing import DataFrame, Series
# base schema definitions
class InSchema(pa.SchemaModel):
str_col: Series[str] = pa.Field(unique=True, isin=[*"abcd"])
int_col: Series[int]
class OutSchema(InSchema):
float_col: pa.typing.Series[float]
# read and validate data from a parquet file
class InSchemaParquet(InSchema):
class Config:
from_format = "parquet"
# output data as a list of dictionary records
class OutSchemaDict(OutSchema):
class Config:
to_format = "dict"
to_format_kwargs = {"orient": "records"}
@pa.check_types
def transform(df: DataFrame[InSchemaParquet]) -> DataFrame[OutSchemaDict]:
return df.assign(float_col=1.1)
The transform function can then take a filepath or buffer containing a parquet file that pandera automatically reads and validates:
import io
import json
buffer = io.BytesIO()
data = pd.DataFrame({"str_col": [*"abc"], "int_col": range(3)})
data.to_parquet(buffer)
buffer.seek(0)
dict_output = transform(buffer)
print(json.dumps(dict_output, indent=4))
Output:
[
{
"str_col": "a",
"int_col": 0,
"float_col": 1.1
},
{
"str_col": "b",
"int_col": 1,
"float_col": 1.1
},
{
"str_col": "c",
"int_col": 2,
"float_col": 1.1
}
]
DataFrameSchemas can now validate geopandas.GeoDataFrame and GeoSeries objects:
import geopandas as gpd
import pandas as pd
import pandera as pa
from shapely.geometry import Polygon
geo_schema = pa.DataFrameSchema({
"geometry": pa.Column("geometry"),
"region": pa.Column(str),
})
geo_df = gpd.GeoDataFrame({
"geometry": [
Polygon(((0, 0), (0, 1), (1, 1), (1, 0))),
Polygon(((0, 0), (0, -1), (-1, -1), (-1, 0)))
],
"region": ["NA", "SA"]
})
geo_schema.validate(geo_df)
You can also define SchemaModel classes with a GeoSeries field type annotation to create validated GeoDataFrames, or use then in @pa.check_types-decorated functions for input/output validation:
from pandera.typing import Series
from pandera.typing.geopandas import GeoDataFrame, GeoSeries
class Schema(pa.SchemaModel):
geometry: GeoSeries
region: Series[str]
# create a geodataframe that's validated on object initialization
df = GeoDataFrame[Schema](
{
'geometry': [
Polygon(((0, 0), (0, 1), (1, 1), (1, 0))),
Polygon(((0, 0), (0, -1), (-1, -1), (-1, 0)))
],
'region': ['NA','SA']
}
)
@pa.dataframe_check works correctly on pandas==1.1.5 (#735)Big shout out to the following folks for your contributions on this release 🎉🎉🎉
add __all__ declaration to root module for better editor autocompletion 42e60c63dddb38b58c6014a14b6fc97b6a3a1e0c
__all__ declaration to root module for better editor autocompletion 42e60c63dddb38b58c6014a14b6fc97b6a3a1e0cBig shout out to the following folks for your contributions on this release 🎉🎉🎉
Pandera now has a discord community! Join us if you need help, want to discuss features/bugs, or help other community members 🤝
Pandera now has a discord community! Join us if you need help, want to discuss features/bugs, or help other community members 🤝
Excited to announce that 0.8.0 is the first release that adds built-in support for additional dataframe types beyond Pandas: you can now use the exact same DataFrameSchema objects or SchemaModel classes to validate Dask, Modin, and Koalas dataframes.
import dask.dataframe as dd
import pandas as pd
import pandera as pa
from pandera.typing import dask, koalas, modin
class Schema(pa.SchemaModel):
state: Series[str]
city: Series[str]
price: Series[int] = pa.Field(in_range={"min_value": 5, "max_value": 20})
@pa.check_types
def dask_function(ddf: dask.DataFrame[Schema]) -> dask.DataFrame[Schema]:
return ddf[ddf["state"] == "CA"]
@pa.check_types
def koalas_function(df: koalas.DataFrame[Schema]) -> koalas.DataFrame[Schema]:
return df[df["state"] == "CA"]
@pa.check_types
def modin_function(df: modin.DataFrame[Schema]) -> modin.DataFrame[Schema]:
return df[df["state"] == "CA"]
And DataFramaSchema objects will work on all dataframe types:
schema: pa.DataFrameSchema = Schema.to_schema()
schema(dask_df)
schema(modin_df)
schema(koalas_df)
pandera.SchemaModels are fully compatible with pydantic:
import pandas as pd
import pandera as pa
from pandera.typing import DataFrame, Series
import pydantic
class SimpleSchema(pa.SchemaModel):
str_col: Series[str] = pa.Field(unique=True)
class PydanticModel(pydantic.BaseModel):
x: int
df: DataFrame[SimpleSchema]
valid_df = pd.DataFrame({"str_col": ["hello", "world"]})
PydanticModel(x=1, df=valid_df)
invalid_df = pd.DataFrame({"str_col": ["hello", "hello"]})
PydanticModel(x=1, df=invalid_df)
Error:
Traceback (most recent call last):
...
ValidationError: 1 validation error for PydanticModel
df
series 'str_col' contains duplicate values:
1 hello
Name: str_col, dtype: object (type=value_error)
Pandera now supports static type-linting of DataFrame types with mypy out of the box so you can catch certain classes of errors at lint-time.
import pandera as pa
from pandera.typing import DataFrame, Series
class Schema(pa.SchemaModel):
id: Series[int]
name: Series[str]
class SchemaOut(pa.SchemaModel):
age: Series[int]
class AnotherSchema(pa.SchemaModel):
foo: Series[int]
def fn(df: DataFrame[Schema]) -> DataFrame[SchemaOut]:
return df.assign(age=30).pipe(DataFrame[SchemaOut]) # mypy okay
def fn_pipe_incorrect_type(df: DataFrame[Schema]) -> DataFrame[SchemaOut]:
return df.assign(age=30).pipe(DataFrame[AnotherSchema]) # mypy error
# error: Argument 1 to "pipe" of "NDFrame" has incompatible type "Type[DataFrame[Any]]";
# expected "Union[Callable[..., DataFrame[SchemaOut]], Tuple[Callable[..., DataFrame[SchemaOut]], str]]" [arg-type] # noqa
schema_df = DataFrame[Schema]({"id": [1], "name": ["foo"]})
pandas_df = pd.DataFrame({"id": [1], "name": ["foo"]})
fn(schema_df) # mypy okay
fn(pandas_df) # mypy error
# error: Argument 1 to "fn" has incompatible type "pandas.core.frame.DataFrame";
# expected "pandera.typing.pandas.DataFrame[Schema]" [arg-type]
Big shout out to the following folks for your contributions on this release 🎉🎉🎉
Strategies should not rely on pandas dtype aliases
unique keyword arg: replace and deprecate allow_duplicates
unique keyword arg: replace and deprecate allow_duplicates (#580)typing.DataFrame class definitions (#576)🎉🎉 Big shout out to all the contributors on this release 🎉🎉
Add support for frictionless schemas (#454) [docs]
pandera.error_formatters.reshape_failure_cases (#560)Big shout out to ✨ @mattHawthorn, @vinisalazar, @cristianmatache, @TColl, @jeffzi, @admackin, and @benkeesey ✨ for your contributions on this release 🎉🎉🎉
Raise error if check_obj.index is MultiIndex when using pandera.Index
Thanks to @jekwatt @cristianmatache @lkadin for your first-time contributions! 🎉🎉🎉
Allow attaching registered dataframe checks by using Config field names
add new method SchemaModel.to_yaml to serialize SchemaModels to yaml #428
SchemaModel.to_yaml to serialize SchemaModels to yaml #428Add SchemaModel column name access through class attributes (#388) @jespercodes @jeffzi 🎉
This release contains two bugfixes:
This release contains two bugfixes:
…schema errors and reason codes. This is a breaking change, but is a minor part of the API and is fairly straightforward to fix (#360).
🎉🎉🎉 Thanks to @jeffzi, @ktroutman, @m1so for your contributions! 🎉🎉🎉
SchemaModel (#329)reset_index, set_index method to DataFrameSchema (#319)SchemaErrors.schema_errors has been changed to failure_cases, and the schema_errors attribute now contains a list of dicts containing schema errors and reason codes. This is a breaking change, but is a minor part of the API and is fairly straightforward to fix (#360).flynt to pre-commit hooks (#325)pandera relied on the packaging package to get version information to determine pandas legacy status. This was an implicit sub-dependency of one of pa
pandera relied on the packaging package to get version information to determine pandas legacy status. This was an implicit sub-dependency of one of pandera's dependencies, which was apparently dropped and led to a bug: #335. This bugfix version explicitly adds packaging.
Deprecate transformers argument in DataFrameSchema init 89c3c9162df8cd3ea0eba4b25e5f6700108c3532
inplace=False argument to schema.validate method to prevent mutation of original dataframe 586ebf3ba550f204f7433247e8628507bb35597c.[hypothesis], [io], [all] available c4716a0ff1f054b8ab0ef5de68e40b5f9de3472d. Thanks @amitripshtos and @jeffzi 🎉check_io decorator for check inputs and outputs of a function 913cbd77345035d9c14d34c02efe9f899aa08e79transformers argument in DataFrameSchema init 89c3c9162df8cd3ea0eba4b25e5f6700108c3532improve failure case reporting more intuitive #232
check_output to the CheckResult namedtuple #251SeriesSchema index specification #270Column object #256check_input decorator when df passed in kwargs #257 thanks @vshulyakDataFrameSchema provides rename_columns method #226 @baskervilski
DataFrameSchema provides rename_columns method #226 @baskervilskicheck_input/output decorators #228lazy validation handles check returning scalar False value 3bf8e72fa2615f086406c1210b893cabe6b232da,
package uses version.py file for single source of truth of package version.
package uses version.py file for single source of truth of package version.
ignore_na keyword argument to Check class 180072e62270f9d6a01819a160d98afff18a4ded drops null columns within the check function before passing to chec
ignore_na keyword argument to Check class 180072e62270f9d6a01819a160d98afff18a4ded
drops null columns within the check function before passing to check_fn. The
SeriesSchemaBase.validate method no longer does this.Column shouldn't need samples arg
0b0d6cc2a73e113c3d82419f525c851cf9e2897blazy kwarg to check_input and check_output 923a197f046c45ee221cdcc4c28d9f8c7c175000schema yaml and script serialization #203
add support for pandas extension types, tested only for pandas >= 1.0.0
PandasDtype enum classNothing published for this version
Schema transformations to add/remove columns from a dataframe schema dcc40f4a3813f4be5f837894a1520a1b02decf84
dtype property to DataFrameSchema and SeriesSchema objects 7aecf163ae3b24e0829a1a7464c870d745495fb7Nothing published for this version
Nothing published for this version
Nothing published for this version
bugfix on the Category datatype and string coercion logic, add more unit tests on dtypes
bugfix on the Category datatype and string coercion logic, add more unit tests on dtypes
improvements in documentation, more comprehensive typing
This release adds docstring examples for all public-facing modules.
This release adds docstring examples for all public-facing modules.
Nothing published for this version
this release drops support for python 2.7: 4d38a4db3eefc13a940b6e47f848583df6f1979d
cosmicBboy -> pandera-dev: e6782f1cb213ed3135fb30768c8edf1c6fa8c536Added support for MultiIndex column and and index validation
MultiIndex column and and index validationDataFrameSchema can validate head, tail, or a random sample of dataframeChecks and Hypothesis checks now support dataframe-level (wide) data validationAdd testing support for python 3.7, improved documentation
Add testing support for python 3.7, improved documentation
This release adds a few nifty features to pandera, special thanks to @mastersplinter and @ralbertazzi:
This release adds a few nifty features to pandera, special thanks to @mastersplinter and @ralbertazzi:
Check class now has a groupby argument, which enables the user to assert properties on subsets of the Column of interest. This opens up the possibility to compare the values or aggregates of values of subsets of a column #42.Hypothesis class, which is a subclass of the Check class. This enables the user to run hypothesis tests on their dataframe as part of a DataFrameSchema definition. Refer to the documentation for more info #43.Columns now have a required argument (default = True), where required=False means that the column is optional #23.SeriesSchemaBase now has an allow_duplicates argument (default = True) #24check_input and check_output decorators 902f1990c3f0ebc07074faee3f3c7e123600a872DataFrameSchema(..., strict=True) means that all columns in the dataframe need to be specified in the schema columns. #34Nothing published for this version
This release adds two new features to pandera.
This release adds two new features to pandera.
Now failure cases in column checks are displayed in a much more compact format, where the failure cases, the index of the dataframe where those failures occur, and the count of failure cases are shown to the user, e.g.
# failure cases:
# index count
# failure_case
# foo1 [0] 1
# foo2 [1] 1
# foo3 [2] 1
DataFrameSchema and ColumnNow the user can coerce the dataframe when calling schema.validate so that
the columns are cast into the expected data-type before performing Checks.
Major change: This release updates changes the API of the DataFrameSchema object. Instead of passing a list of Columns, you now pass a dictionary wher
DataFrameSchema object.
Instead of passing a list of Columns, you now pass a dictionary where the keys are column_names
values are Column objects. This makes the API feel a lot more familiar for pandas users, who may
often define DataFrames in a similar way (see README for details).Validator to Check for brevity and clarity (accordingly renamed validator_{input, output}
to check_{input, output}.PandasDtype so they can be accessed in pandera namespace:
Bool, DatetTime, Category, Float, Int, Object, String, TimedeltaYour coding agent can read these notes before it upgrades. Set up the MCP server →