NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2010 most downloaded on PyPI
Pinecone Python SDK
Last release 1 months ago
03 Sep 2026
Release timing varies
gaps range from 8 days to 3 months
Nearly every release is documented
notes for 41 of 43 stable releases
1 version withdrawn
withdrawn after publishing
2 years old
97 releases · first in 2024
h2 (Python) 4.3.0 to 4.4.1 , for CVE-2026-71554 . It arrives with the default install, as a transitive of httpx[http2] . 4.4.1 also rejects duplicate…
Version 10.0.0 is the SDK for the Pinecone 2026-07 API. It runs on Python 3.10 through 3.14 and sends 2026-07 on every plane by default. The release graduates the schema-based document interface, SchemaBuilder, and the rest of the 2026-07 API out of pinecone.preview onto the main client objects. There are breaking changes: a short list of the most important changes are outlined below, and Migrating to V10 is the complete accounting.
If your workload upserts and queries an existing index, upgrading does not change what your code means. upsert, query, fetch, update, delete, list, and describe_index_stats take the same arguments and return the same shapes they did on 9.x; pc.Index("movies") and pc.index("movies") both still hand you an index, though only pc.index() narrows its return type, so reach index.documents through that one; and none of those operations is deprecated or scheduled for removal. Indexes you created before 2026-07 go on being served by them. Three things do change on that path, each covered below: a batched upsert paces itself differently, some invalid calls now raise locally instead of being refused by the server, and gRPC upsert_from_dataframe collects partial failures rather than raising on the first one.
The breaking changes are concentrated in pc.indexes.create's arguments and return model, and in the surfaces that graduated out of pinecone.preview. Flat create_index(...) with dimension=, metric=, or spec= needs no edits: it takes its 9.1.0 parameters in the 9.1.0 order, positionally or by keyword. Which family of operations serves an index follows the index itself, not the SDK version — an index created before 2026-07 is addressed through the vector operations above, and one created with a 2026-07 document schema through index.documents.
Documents API. Records are addressed as documents: a JSON object with an _id, the fields you declared in the index schema, and any metadata alongside them. Upsert, search, fetch, list, update, and delete all live on index.documents, on the sync and async clients alike, and index.documents.batch_upsert handles large loads with the same admission gate and partial-failure reporting the vector path uses. See Quickstart.
from pinecone import DenseVectorQuery, Pinecone
index = Pinecone().index("quickstart")
index.documents.upsert(
namespace="movies",
documents=[
{
"_id": "movie-001",
"embedding": [0.1, 0.2, 0.3],
"title": "Arrival",
},
],
)
results = index.documents.search(
namespace="movies",
top_k=3,
score_by=[DenseVectorQuery(field="embedding", values=[0.1, 0.2, 0.3])],
)SchemaBuilder. An index's fields are declared as a schema, and SchemaBuilder builds that schema in Python instead of by hand-assembling nested dicts. Every add_* method returns the builder, and build() returns a plain {"fields": {...}} dict you hand straight to pc.indexes.create(schema=...). See Schema Builder.
from pinecone import Pinecone, SchemaBuilder
schema = (
SchemaBuilder()
.add_dense_vector_field("embedding", dimension=1024, metric="cosine")
.add_string_field("title", full_text_search={"language": "en"})
.build()
)
pc = Pinecone()
pc.indexes.create(
name="product-search",
schema=schema,
deployment={
"deployment_type": "managed",
"cloud": "aws",
"region": "us-east-1",
},
)Full-text search configuration. String fields carry their own full-text search config: language, stemming, stop words, and character n-grams for substring and autocomplete matching, set per field at index creation. NgramConfig and FullTextSearchConfig are exported from pinecone for building the config as typed objects; n-grams cannot be combined with stemming or stop words. See Schema Builder.
schema = (
SchemaBuilder()
.add_string_field(
"title",
full_text_search={"ngram": {"min_gram": 2, "max_gram": 4}},
)
.build()
)Backup schedules. Attach a recurring cadence to an index and backups keep happening without a caller. Schedules live on pc.backup_schedules, with create, list, describe, update, delete, and a per-schedule run history. frequency takes exactly daily, weekly, or monthly; there is no cron support, and only one enabled schedule per index is allowed. See Backups and restore.
schedule = pc.backup_schedules.create(
index_name="product-search",
name="daily-backup",
frequency="daily",
retention_days=90,
)
for run in pc.backup_schedules.iter_history(schedule_id=schedule.schedule_id):
print(run.backup_id, run.status, run.record_count)Organization management on the Admin client. Admin was OAuth, organizations, projects, and API keys. It now also manages the people and machines in an organization: admin.users, admin.invites, admin.service_accounts, and admin.role_bindings. Role bindings are the whole authorization model: one role, one principal, one scope, and nothing else confers permissions. See Admin.
from pinecone.admin import Admin
admin = Admin(client_id="...", client_secret="...")
admin.invites.create(
email="newhire@acme.com",
role_bindings=[
{"resource_type": "organization", "role": "OrgMember"},
],
)
created = admin.service_accounts.create(
name="ci-deploy",
role_bindings=[
{
"resource_type": "project",
"resource_id": project_id,
"role": "ProjectEditor",
},
],
)
print(created.client_secret) # returned exactly onceAssistant operations API. Assistant file writes are asynchronous server-side, and the operations behind them are now first-class and can be listed and described, so a job started with timeout=-1 can be followed rather than inferred from file status. Successes and failures are both kept for 30 days. See Assistant.
operation = pc.assistants.describe_operation(
assistant_name="my-assistant", operation_id="op-1234-abcd-5678"
)
print(operation.status, operation.percent_complete, operation.error)
stuck = pc.assistants.list_operations(
assistant_name="my-assistant",
operation_type="upload_file",
status="Processing",
).to_list()Dedicated read capacity. An index can be provisioned on dedicated read nodes instead of on-demand capacity, at create time, on configure, and on a restore. One top-level read_capacity= argument covers managed and BYOC indexes, and it reads back from IndexModel.read_capacity. pc.create_index_from_backup(..., read_capacity=...) applies the same configuration to a restored index in one call. See Backups and restore.
pc.indexes.create(
name="product-search",
schema=schema,
deployment={
"deployment_type": "managed",
"cloud": "aws",
"region": "us-east-1",
},
read_capacity={
"mode": "Dedicated",
"dedicated": {
"node_type": "t1",
"scaling": "Manual",
"manual": {"shards": 2, "replicas": 2},
},
},
)Bulk ingest deadlines and concurrency. Batched upserts take a total_timeout, a deadline for the whole call rather than for one attempt of one batch. When it expires the SDK stops submitting; batches in flight finish, and everything unsent comes back as failed items to retry. max_concurrency defaults to 8 and is capped by an adaptive per-host admission gate, so a backend under pressure gets backpressure instead of every batch at once. Identical on REST sync, asyncio, and gRPC. See How Bulk Ingest Behaves.
response = index.upsert(
vectors=vectors,
batch_size=200,
max_concurrency=16,
total_timeout=1800,
)
# Batch counters and errors carry information only when batch_size is set.
print(response.upserted_count, response.failed_item_count)
print([(e.disposition, e.retryable) for e in response.errors])upsert_from_dataframe built for large ingests. The DataFrame path has the same signature on REST, asyncio, and gRPC, and it now gets the full bulk-ingest toolkit: a per-batch timeout, a max_concurrency ceiling, a total_timeout deadline for the whole job, and on_error to choose between collecting partial failures into the response or raising. max_concurrency is unset by default here, unlike upsert, which defaults to 8. Rows that did not land come back in response.failed_items, ready to wrap in a new DataFrame and feed back in. See Reliable Large Ingests with upsert_from_dataframe.
response = index.upsert_from_dataframe(
df,
namespace="movies",
batch_size=500,
max_concurrency=8,
total_timeout=1800,
)
if response.failed_items:
leftovers = pd.DataFrame(response.failed_items)
index.upsert_from_dataframe(leftovers, namespace="movies", on_error="raise")Python 3.14. The SDK is tested and classified on 3.10 through 3.14; the requires-python floor stays at 3.10.
This is the headline list, not the complete one. Read Migrating to V10 for every change, including the ones that affect only narrow call shapes.
pinecone.preview is removed with no shim. Preview, AsyncPreview, and every preview.* module are gone. Each call site maps onto the main client:
| 9.x | 10.0.0 |
|---|---|
pc.preview.indexes |
pc.indexes |
pc.preview.index(name=...) / (host=...) |
pc.index(name=...) / (host=...), or pc.index("my-index") |
pc.preview.index(...).documents.upsert(...) |
pc.index(...).documents.upsert(...) |
pc.preview.close() |
pc.close() |
from pinecone.preview import SchemaBuilder |
from pinecone import SchemaBuilder |
AsyncPinecone.index() also became a coroutine — AsyncPreview.index() resolved its host lazily, so it needs an await now.
IndexModel is built from schema and deployment. .created_at is gone. .spec, .embed, .dimension, .metric, and .vector_type still work as deprecated views computed from the new fields, including through index["..."] and to_dict(), but they raise when the schema gives them no single field to resolve to — more than one vector field, or none — and the message names the fields it found.
Bulk ingest defaults changed. max_concurrency went from 4 to 8 on every upsert path, an adaptive admission gate can hold batches back or abandon the rest of a call against a dead backend, and total_timeout bounds the whole call. Pass max_concurrency=4 for the old cap, and inspect response.errors rather than assuming a call that returned did all the work.
GrpcIndex.upsert_from_dataframe no longer raises on partial failure. Failures are collected into the response, so an existing except block around it goes dead. Pass on_error="raise" to restore the previous behavior. Partial failures on gRPC upsert_from_dataframe covers what the response carries and how to retry from it.
pc.indexes.list() returns a Paginator[IndexModel], not an IndexList, so .names() is not available on that path. pc.list_indexes() is unchanged.
Client-side validation raises before the request is sent. top_k outside 1-10000, a page limit out of range, an id over 512 characters, an empty filter={}, mutually exclusive arguments, and malformed namespace names now raise PineconeValueError or PineconeTypeError locally, where 9.x let the server refuse them. Neither is an ApiError subclass, so an except ApiError handler will not see them; except PineconeError covers both.
Assistant file progress fields moved to the operations API. AssistantFileModel.percent_done and .error_message are removed; call pc.assistants.describe_operation(...) and read status, percent_complete, and error.
BackupModel.dimension and .metric are removed. Read .schema instead.
The default API version header is now 2026-07 on every plane. To pin an older version, set it yourself: additional_headers={"X-Pinecone-Api-Version": "2026-01"}.
These paths still work and are not scheduled for removal in 10.x.
dimension=, metric=, vector_type=, and spec= on create_index and pc.indexes.create. Pass schema= and deployment= instead. replicas=, pod_type=, and serverless_read_capacity= on configure are deprecated the same way, translated into deployment= and read_capacity=.IndexModel.spec, .embed, .dimension, .metric, and .vector_type, along with the IndexSpec, ServerlessSpecInfo, PodSpecInfo, ByocSpecInfo, and ModelIndexEmbed classes. Read .deployment and .schema.fields instead; that is where the server reports them and the only place they appear for an index with several vector fields.pool_threads on the client and async_req=True on data-plane methods. Use AsyncPinecone or a ThreadPoolExecutor; these exist for backcompat only.RerankModel.Pinecone_Rerank_V0 is still a member but deprecated, and most projects now get a permission error when they ask for it.None of these emit a DeprecationWarning yet, so -W error::DeprecationWarning will not find your call sites. Grep for them.
Two new members of the exception tree apply across every surface: PaymentRequiredError (402) and FailedPreconditionError (412), both ApiError subclasses.
Client and connection
ssl_ca_certs, ssl_verify, and proxy_url reach the transport that opens the socket, so they take effect.pc.index(name) is overloaded on grpc=, so index.documents type-checks under mypy and pyright.pc.preview raises an AttributeError that names its replacement and links the migration guide, rather than Python's bare message.Indexes and namespaces
create_index, configure_index, and list_indexes still work on top of the schema and deployment API.NamespaceDescription.size_bytes is populated.Inference
model is an EmbedModel member.EmbedModel members for the models available at 2026-07.top_n >= 1 is enforced on the async rerank path, and argument conflicts are reported before enum validation.Assistant
timeout= the read floor is 300 seconds.content_filter_results, and StreamMessageStart carries an id.Packaging and types
TokenResponse are exported at the top level.anyio>=4.0 is declared as an explicit runtime dependency instead of relied on through httpx.LICENSE file into the site-packages root; it ships in the sdist and in the wheel's metadata.[dev] extra is removed; use uv sync --group dev.Four fixes in the SDK's own code:
/, or a . or .. segment, previously survived into the request path and addressed a different endpoint instead of failing: describe_namespace(name=".") requested the namespace list route, and describe_namespace(name="..") requested /. Encoding now happens in HTTPClient and AsyncHTTPClient rather than at each call site, so it covers every path the SDK builds. The exposure was bounded to misdirected reads: none of the collapsing paths resolved to a DELETE, PUT, or PATCH route.TokenResponse.__repr__ masks the access token. It previously rendered the live Bearer token in full, so any crash reporter that captures frame locals — Sentry's Python SDK does by default — wrote the credential verbatim into the error report. repr() and str() now keep only the last four characters. The masking stops there: to_dict() and JSON encoding still return the token in full, so do not serialize the object wholesale into a log line or a cache.Admin no longer keeps the raw client_secret where a repr() reaches it. The token minter held the secret as a bound argument of a functools.partial, whose repr() renders the arguments it closed over.PINECONE_GRPC_SCHEME=http against a host that is neither loopback nor RFC 1918 private, the SDK now emits a RuntimeWarning, because the API key crosses a public network in the clear.And the dependency tree. As of 10.0.0 no advisory against the SDK's dependencies is outstanding. Three of those upgrades reach code that ships to you:
h2 (Rust) 0.4.13 to 0.4.19, for RUSTSEC-2026-0258: the HTTP/2 codec queued empty DATA frames without bound, so a hostile or compromised endpoint could drive unbounded memory growth or a length-overflow panic. It reached the shipped gRPC extension as a runtime transitive, through hyper and tonic, and 9.1.0 carried the same version. If you use the gRPC transport, upgrade for this alone.h2 (Python) 4.3.0 to 4.4.1, for CVE-2026-71554. It arrives with the default install, as a transitive of httpx[http2]. 4.4.1 also rejects duplicate Host headers and conflicting content-length values rather than accepting the first of each.pyo3 0.24.1 to 0.29.2, for GHSA-36hh-v3qg-5jq4 and GHSA-chgr-c6px-7xpp, compiled into the shipped extension. Neither advisory is reachable from this codebase, since the affected APIs have no call sites here; the upgrade went in regardless.The remaining advisories were against development and documentation tooling — soupsieve and tornado, reached through beautifulsoup4 and ipykernel — and have never been present in a published wheel.
The Rust dependency tree is audited on every push, so an advisory against a crate inside the gRPC extension surfaces on the commit that introduces it, and the PyPI publish step is pinned to a commit SHA.
pip install --upgrade pinecone
# or
uv add 'pinecone>=10,<11'Then read Migrating to V10.
One column per month.
Nothing published for this version
v9.1.0 is the retry-resilience release. It overhauls how the SDK behaves under throttling: decorrelated jitter on every transport, plus automatic per-
v9.1.0 is the retry-resilience release. It overhauls how the SDK behaves under throttling: decorrelated jitter on every transport, plus automatic per-host concurrency back-off that shrinks in-flight bulk requests during a 429 storm and recovers as pressure eases. A new RateLimitError lets callers catch throttling distinctly, and a dedicated retries guide documents the full model. Secondary changes round out GrpcIndex parity with the REST clients and restore the v8 search(query={...}) request shape.
AsyncIndex.query_namespaces no longer fires one unbounded request per namespace — a common cause of self-inflicted 429 storms. It's now capped at 10 in-flight.RateLimitError for HTTP 429. Rate-limit responses raise a dedicated pinecone.errors.RateLimitError, so you can catch throttling distinctly from other server errors.docs/guides/retries.md — defaults, jitter math, the adaptive-concurrency ceiling, and the limits of in-process retries under multi-process / serverless deployment.GrpcIndex parity with the REST clients. query_namespaces, query_namespaces_async, fetch_by_metadata, and the full bulk-import API now exist on GrpcIndex, so switching to grpc=True for throughput no longer loses functionality.search(query={...}) syntax from v8 has been restored. Thanks to @pragnyanramtha, the single-argument request-body form that v9.0.x rejected works again on Index, AsyncIndex, and GrpcIndex.REST and gRPC retries now use decorrelated jitter — each delay is drawn from a window that grows with the previous delay, capped at max_wait. This replaces the previous fixed full-jitter backoff and spreads retries from many clients across time instead of bunching them at the same instant. See the retries guide for the exact formula and defaults.
Bulk operations now self-tune their concurrency under throttling. When the backend rate-limits a host (HTTP 429/503, or gRPC RESOURCE_EXHAUSTED), the SDK reduces the number of in-flight requests it allows to that host, then recovers after a streak of successes — the same AIMD (additive-increase / multiplicative-decrease) approach used by TCP congestion control. It's per host, automatic, and requires no configuration; max_concurrency remains the ceiling the SDK tunes beneath.
RateLimitErrorHTTP 429 responses now raise RateLimitError, a subclass of ApiError, so throttling can be caught distinctly from other server errors:
from pinecone.errors import RateLimitError
try:
pc.indexes.describe("my-index")
except RateLimitError:
# back off and retry, or surface to your orchestrator
...
Exported from pinecone.errors and the top-level pinecone namespace; a RateLimitException alias matches the existing exception-naming pattern.
The retry and concurrency layers emit namespaced log records with consistent key=value fields, so you can diagnose throttling from logs without adding instrumentation:
pinecone._internal.http_client — DEBUG on each throttled retry attempt (status, host, attempt N/total, computed delay).pinecone._internal.adaptive — DEBUG on each concurrency adjustment; INFO the first time a host is throttled, with an actionable message.Enable these by setting the log level on the namespaces above. See the retries guide for examples.
These affect callers who set explicit RetryConfig values. Default-config callers get the new behavior transparently.
RetryConfig.max_retries on REST now counts retries, not total attemptsPreviously the REST transport treated max_retries as the total attempt count, while gRPC counted retries only — so the same RetryConfig(max_retries=N) behaved differently across transports. REST now matches gRPC: max_retries=N means N retries after the initial attempt (up to N+1 total requests). A caller passing max_retries=3 will now see 4 total requests on a persistent failure instead of 3.
RetryConfig.backoff_factor default: 2.0 → 0.25backoff_factor is no longer an exponential multiplier — it's the minimum delay floor in the decorrelated-jitter formula. The default dropped from 2.0 to 0.25, giving a first-retry mean of ~0.5s — substantially shorter than v8, whose 2.0 default produced roughly 4× longer waits. Callers who set backoff_factor explicitly should re-read its docstring and the migration note in the retries guide; a v8 backoff_factor=2.0 maps to roughly 0.5 under the new semantics for a comparable first-retry delay.
GrpcIndex parityGrpcIndex was missing surface area that Index and AsyncIndex already had in v9.0.x. Switching to grpc=True for throughput should not silently lose functionality.
query_namespaces and query_namespaces_asyncMulti-namespace fan-out queries, with the same validation and aggregation semantics as the REST path.
grpc_index = pc.index("my-index", grpc=True)
results = grpc_index.query_namespaces(
vector=[0.1, 0.2, ...],
namespaces=["a", "b", "c"],
top_k=10,
metric="cosine",
)
fetch_by_metadataSame signature as Index.fetch_by_metadata.
start_import, describe_import, cancel_import, list_imports, and list_imports_paginated are now available on GrpcIndex.
GrpcIndex.search dense-vector fixGrpcIndex.search() sent bare list[float] queries in the wrong wire shape, so dense-vector queries failed. It now matches Index.search / AsyncIndex.search, and dense-vector queries via GrpcIndex.search() work.
search(query={...}) request body on all surfaces (#668)This release shape was contributed by the community — thanks to @pragnyanramtha for the PR.
Index.search, AsyncIndex.search, and GrpcIndex.search accept a single query= argument carrying the request body (a SearchQuery or a plain Mapping) as an alternative to the flat keyword form — restoring a v8 shape that v9.0.x rejected. The contracts are tight:
query= and direct kwargs (e.g. vector=, top_k=) raises TypeError, listing the conflicting names.query= (e.g. an int) raises a clear TypeError instead of a confusing downstream error.Fixes #661.
docs/guides/retries.md — defaults (REST vs gRPC), RetryConfig fields, the jitter formula with examples, adaptive concurrency, transport differences, and honest limitations under multi-process / serverless deployment. See the rendered docs page at Retries and Resiliencedocs/guides/error-handling.md now cross-references the new guide instead of duplicating it (the ## Retries heading is preserved so existing anchors resolve). See the rendered docs page at Error handlingstart_import docstring error_mode default corrected from "abort" to "continue".create_index_from_backup docstrings now point to backups.create / backups.list as the sources for a backup_id.GrpcIndex.upsert_records namespace docstring no longer suggests "", which validation rejects.Please file issues at https://github.com/pinecone-io/python-sdk/issues — include the SDK version (pinecone.__version__), Python version, and a minimal reproduction. For throttle / retry behavior, attach DEBUG logs from pinecone._internal.http_client and pinecone._internal.adaptive if possible.
Full Changelog: https://github.com/pinecone-io/python-sdk/compare/v9.0.1...v9.1.0
The v9.0.0 field= kwarg continues to work as a deprecated alias and will be removed in a future release. pc.preview is not covered by SemVer; see the…
v9.0.1 is a polish release. It rolls up two weeks of fixes against backend response shapes, new client-side validation, completes the gRPC namespace API, widens public input types for ergonomic flexibility, and lands a deep pass on docstrings and migration guidance.
create_namespace, describe_namespace, delete_namespace, list_namespaces, list_namespaces_paginated), sparse/hybrid vector queries on Index.search / AsyncIndex.search, pc.preview.indexes.describe_backup, pc.indexes.configure(serverless_read_capacity=...), pc.assistants.create(..., environment=...), and several response fields the backend was returning but the SDK was dropping. See Filling v9.0.0 gaps.read_capacity are correctly forwarded everywhere. Several v9.0.0 paths (BYOC, IntegratedSpec, ServerlessSpec, async creation, the create_index shim, configure) silently dropped these fields; all are now end-to-end correct.list[T] now accept any Sequence[T]; dict[str, str] and dict[str, Any] parameters now accept Mapping. Existing call sites are unaffected, but tuple, generator-backed sequences, and Mapping subclasses now type-check and work without conversion.pinecone/__init__.pyi stub fixes type resolution in Spyder, Eric, Cython, and other editors that don't follow from … import * re-exports. mypy and Pyright/Pylance users are unaffected.A handful of methods and fields that should have shipped in v9.0.0 were inadvertently left out. They are restored in v9.0.1.
These methods exist on the REST Index but were missing from GrpcIndex in v9.0.0:
grpc_index = pc.index("my-index", grpc=True)
grpc_index.create_namespace("orders")
ns = grpc_index.describe_namespace("orders")
for ns in grpc_index.list_namespaces(): # auto-paginating generator
print(ns.name)
page = grpc_index.list_namespaces_paginated(limit=50)
grpc_index.delete_namespace("orders")
Index.search / AsyncIndex.search: sparse and hybrid queriesThe REST search surface accepts sparse and hybrid vector queries, but v9.0.0 only plumbed dense list[float] through Index.search / AsyncIndex.search. Sparse and hybrid inputs are now accepted as a dict alongside dense lists:
index.search(
namespace="docs",
query={"vector": {"values": [0.1, 0.2, ...], "indices": [3, 17, 42, ...]}},
top_k=10,
)
Keys are validated and normalized client-side; invalid combinations (e.g. values without indices on sparse) raise ValidationError before the HTTP call.
describe_backuppc.preview.indexes.describe_backup(...) and its async counterpart were omitted from v9.0.0 and are now present.
The backend accepts these parameters on preview index create / configure, but the SDK wasn't forwarding them:
PreviewIndexes.create(..., source_collection=..., source_backup_id=..., cmek_id=...)PreviewIndexes.configure(..., deployment=...) (pod index patching)PreviewIndexModel now exposes private_host, source_collection, source_backup_id, and cmek_id on the response sideenvironment parameterpc.assistants.create(..., environment=...) — the backend accepted it; the SDK was not forwarding it.
pc.indexes.configure(serverless_read_capacity=...) — the parameter existed on the REST surface but was not plumbed through the SDK.IndexModel.private_host — surfaces private-endpoint hosts the backend was already returning.ByocSpecInfo.schema — was being dropped during response decoding.The public surface now accepts the broadest reasonable type at every input parameter. This is purely an extension — existing list / dict call sites continue to work — but it removes a class of "I have to convert to a list first" friction:
list[T] → Sequence[T] on input parameters (DX-0140), including embed, rerank, and assistant messages/roles (pc.assistants.chat, etc.)dict[str, str] → Mapping[str, str] for headers and tagsdict[str, Any] → Mapping[str, Any] for read-only params on Index operations (DX-0141)This also unblocks downstream type-checked callers that already had widened parameter types and were forced into cast() or list(...) conversions.
pinecone/__init__.pyi stub for Jedi-based IDEs (DX-0145)Editors that use the Jedi static analyzer (Spyder, Eric, Cython, some older VS Code setups) don't traverse from … import * re-exports, so from pinecone import Pinecone showed no autocomplete in those environments. A new pinecone/__init__.pyi stub enumerates the public surface explicitly. mypy, Pyright, and modern Pylance users are unaffected (they already resolved through the runtime module); Jedi users now see full autocomplete and go-to-definition.
A CI typecheck-drift gate (DX-0146) ensures the stub stays in sync with the runtime pinecone/__init__.py.
PreviewTextQuery: field renamed to fieldsPreviewTextQuery.field is renamed to fields so it can accept multiple field names for multi-field text search. The v9.0.0 field= kwarg continues to work as a deprecated alias and will be removed in a future release. pc.preview is not covered by SemVer; see the v9.0.0 release notes.
These are not breaking under SemVer (they reject inputs the backend was already rejecting), but a few new client-side checks may surface as ValidationError where v9.0.0 surfaced as a server-side error. If your code wraps these calls in try/except, no changes are needed — ValidationError is a subclass of PineconeException — but the error class and message will differ.
top_k > 10000 on GrpcIndex.query is now rejected client-side. REST had this check already.top_n < 1 on pc.inference.rerank(...) is now rejected client-side.limit < 1 on Index.fetch_by_metadata is now rejected client-side (AGT-0044).upsert_records rejects records with an invalid _id field (non-string, empty, or containing null bytes) before the HTTP call.pc.collections.create(name=...) validates the resource name client-side (length, character set).Projects.create/update validates the project name (length ≤ 50, no null bytes) and refuses negative max_pods.ApiKeys.create enforces the 80-character name length cap.PreviewDocuments.fetch(ids=...) rejects empty ids lists.Two PreviewDocuments parameters that were never functional were removed:
PreviewDocuments.fetch(..., filter=...) — the filter parameter was dropped; it had no effect on the backend.PreviewDocuments.delete(..., filter=...) — same.If you were passing filter= to either of these on v9.0.0, the call will now raise TypeError. The argument was never being sent, so removing it does not change observable behavior.
fix(grpc): Cancelled + "Timeout expired" from the gRPC backend is now raised as PineconeTimeoutError instead of bubbling as the lower-level grpc.RpcError (CI-0052). Code that caught PineconeTimeoutError already handles this; code that caught only grpc.RpcError will no longer see those specific timeouts.
A number of v9.0.0 response structs had fields typed as non-optional that the backend in fact returns as null in some states, causing decode failures. These have been widened:
IndexModel.host — optional (null during creation / Terminating states)BackupModel.tags — dict[str, Any] (was dict[str, str])BackupModel.schema — real schema field (was an always-None stub)BackupModel.metric — removed (never returned by the backend)PodSpecInfo.replicas, .shards, .pods — optionalRestoreJobModel.created_at — optionalAPIKeyModel.name — str | None; the dead description field is removedStreamMessageEnd.usage — optional (handles null from streaming responses)PreviewBackupModel.tags — dict[str, Any] (was dict[str, str])PreviewIndexModel.host — optionalPreviewManagedDeployment.environment — optionalPreviewReadCapacityStatus.error_message — new field for failure statesModelInfoSupportedParameter.min / .max — int | float | None (v9.0.0 declared float, but the backend returns ints for some parameters)ByocSpecInfo.schema — newly surfacedIf you were destructuring these fields and bypassing None handling, add a guard. mypy will surface the new None possibilities under --strict.
schema and read_capacity through every index-creation path: integrated (pc.indexes.create_for_model), BYOC (build_byoc_body), ServerlessSpec, IntegratedSpec, the async creation path, and the create_index backcompat shim. On v9.0.0 these fields were dropped at the boundary depending on which path you called.pc.indexes.configure(...): stop pre-merging tags client-side; the backend handles tag merges, and the client-side step was overwriting partial updates. Same job, fewer bugs.pc.indexes.configure(...): relax validate_read_capacity to allow partial Dedicated patches (e.g. changing min_replicas without re-specifying max_replicas).configure boundary so that schema={"fields": {...}} and the shorter schema={"my_field": ...} are both accepted.IndexTerminatedError when an index enters Terminating or Disabled during a poll, instead of looping until the operator's timeout.EmbedConfig: include dimension in the create request (was silently dropped).pc.assistants.create(...): omit metadata from the request body when not provided. v9.0.0 sent metadata: null, which a stricter backend rejected.pc.assistants.list_page / list_files_page: upgrade to API version v202604, fix query-parameter names and pagination-token parsing, add page_size.metadata field as multipart form data (was being sent as a query param, which silently dropped the value).ChatCompletionMessage: role and content are now optional (matches the streaming assistant-role chunk shape where these arrive on a later chunk).ChatCompletionStreamChunk: expose the usage field, which had been dropped during struct decoding.EntailmentResult.reasoning: populate from the evaluate_alignment response (was always empty on v9.0.0).OperationModel.error: map from the JSON field error_message (was empty for failed operations).PINECONE_PLUGIN_ASSISTANT_CONTROL_HOST / PINECONE_PLUGIN_ASSISTANT_DATA_HOST env vars at client construction (legacy plugin compatibility).Terminated and Terminating as terminal states in _poll_until_ready, so creation polling exits promptly when an assistant transitions to a terminal failure state.Index.search / AsyncIndex.search: validate and normalize sparse/hybrid vector dict keys; wrap dense vector queries in {"values": ...} for backend serde compatibility.query_namespaces: preserve input namespace order in the result so ties break deterministically.upsert_records: drop the redundant id key when _id is present (the backend treats either, but sending both caused a 400).pc.backups.create(...): omit description from the request body when not provided.pc.indexes.create_from_backup(timeout=-1): return CreateIndexFromBackupResponse directly instead of attempting to poll with a negative timeout.limit=10 default on Backups.list / RestoreJobs.list so backend-default pagination is used.PreviewIndexes.create and configure: full tag validation (was incomplete on v9.0.0).PreviewIndexes.configure: validate schema fields are semantic_text only.PreviewIndexes.list_backups: forward limit to the server in fetch_page (was hardcoded).PreviewBooleanField and PreviewLegacyIntegerField are now in the schema union; v9.0.0 rejected schemas containing them.Projects.create / update: client-side name length and null-byte validation.Projects.update: reject negative max_pods before the HTTP call.ApiKeys.create: enforce the 80-character name length cap.ApiKeys.create: remove the dead description parameter (server doesn't accept it).rerank(...): validate top_n >= 1 before the HTTP call.list_models(...): reject type=rerank combined with a vector_type argument.ModelInfoSupportedParameter.min/max typing widened to int | float | None.import_data(...): omit errorMode from the wire payload when error_mode is not specified (was sending null, which the backend rejected).schema and read_capacity through async index-creation shims so that v8-style positional calls continue to behave correctly.AssistantModel.context shim to accept a messages= argument (matches the deprecated v8 signature).A broad pass over docstrings, RST formatting, and the v9 migration guide. Highlights:
pc.index() rather than the deprecated shim:meta private: so they no longer clutter generated docsPinecone.index, Pinecone.config, Pinecone.close, Pinecone.assistant, Pinecone.backups, Pinecone.collections, Pinecone.assistants: complete Args, Returns, Raises, and Examples sections where missingupsert_from_dataframe max_concurrency claim corrected in performance guidegrpc extra_AssistantNamespaceProxy; build remains warnings-as-errors cleanquery=, spec_embed=, field=), unsafe .get() calls, deprecated import paths, false immutability claimsThe Sphinx docs build remains green with -W (warnings-as-errors).
pytest-xdist (-n 6 --dist=loadfile); the release-readiness job's timeout was raised from 60 → 90 min as a margin.abi3audit pass runs on every wheel produced by the cross-platform matrix; catches stable-ABI violations before publish (the multi-Python smoke job is the runtime complement).abi3audit and auditwheel in their own venvs so they no longer leak into the verify Python's pip check.pinecone/ tree before running so the installed wheel is actually exercised.dev-version resolution now uses the artifacts API directly rather than walking workflow runs (avoids missing the most recent real publish when intermediate runs are no-op skips).If something worked on v9.0.0 and doesn't on v9.0.1, please open an issue at https://github.com/pinecone-io/python-sdk/issues with the same details requested in the v9.0.0 notes (versions, minimal repro, expected vs actual). Compatibility regressions are prioritized.
Full Changelog: https://github.com/pinecone-io/python-sdk/compare/v9.0.0...v9.0.1
v9.0.1 is a polish release. It rolls up two weeks of fixes against backend response shapes, new client-side validation, completes the gRPC namespace API, widens public input types for ergonomic flexibility, and lands a deep pass on docstrings and migration guidance.
create_namespace, describe_namespace, delete_namespace, list_namespaces, list_namespaces_paginated), sparse/hybrid vector queries on Index.search / AsyncIndex.search, pc.preview.indexes.describe_backup, pc.indexes.configure(serverless_read_capacity=...), pc.assistants.create(..., environment=...), and several response fields the backend was returning but the SDK was dropping. See Filling v9.0.0 gaps.read_capacity are correctly forwarded everywhere. Several v9.0.0 paths (BYOC, IntegratedSpec, ServerlessSpec, async creation, the create_index shim, configure) silently dropped these fields; all are now end-to-end correct.list[T] now accept any Sequence[T]; dict[str, str] and dict[str, Any] parameters now accept Mapping. Existing call sites are unaffected, but tuple, generator-backed sequences, and Mapping subclasses now type-check and work without conversion.pinecone/__init__.pyi stub fixes type resolution in Spyder, Eric, Cython, and other editors that don't follow from … import * re-exports. mypy and Pyright/Pylance users are unaffected.A handful of methods and fields that should have shipped in v9.0.0 were inadvertently left out. They are restored in v9.0.1.
These methods exist on the REST Index but were missing from GrpcIndex in v9.0.0:
grpc_index = pc.index("my-index", grpc=True)
grpc_index.create_namespace("orders")
ns = grpc_index.describe_namespace("orders")
for ns in grpc_index.list_namespaces(): # auto-paginating generator
print(ns.name)
page = grpc_index.list_namespaces_paginated(limit=50)
grpc_index.delete_namespace("orders")Index.search / AsyncIndex.search: sparse and hybrid queriesThe REST search surface accepts sparse and hybrid vector queries, but v9.0.0 only plumbed dense list[float] through Index.search / AsyncIndex.search. Sparse and hybrid inputs are now accepted as a dict alongside dense lists:
index.search(
namespace="docs",
query={"vector": {"values": [0.1, 0.2, ...], "indices": [3, 17, 42, ...]}},
top_k=10,
)Keys are validated and normalized client-side; invalid combinations (e.g. values without indices on sparse) raise ValidationError before the HTTP call.
describe_backuppc.preview.indexes.describe_backup(...) and its async counterpart were omitted from v9.0.0 and are now present.
The backend accepts these parameters on preview index create / configure, but the SDK wasn't forwarding them:
PreviewIndexes.create(..., source_collection=..., source_backup_id=..., cmek_id=...)PreviewIndexes.configure(..., deployment=...) (pod index patching)PreviewIndexModel now exposes private_host, source_collection, source_backup_id, and cmek_id on the response sideenvironment parameterpc.assistants.create(..., environment=...) — the backend accepted it; the SDK was not forwarding it.
pc.indexes.configure(serverless_read_capacity=...) — the parameter existed on the REST surface but was not plumbed through the SDK.IndexModel.private_host — surfaces private-endpoint hosts the backend was already returning.ByocSpecInfo.schema — was being dropped during response decoding.The public surface now accepts the broadest reasonable type at every input parameter. This is purely an extension — existing list / dict call sites continue to work — but it removes a class of "I have to convert to a list first" friction:
list[T] → Sequence[T] on input parameters (DX-0140), including embed, rerank, and assistant messages/roles (pc.assistants.chat, etc.)dict[str, str] → Mapping[str, str] for headers and tagsdict[str, Any] → Mapping[str, Any] for read-only params on Index operations (DX-0141)This also unblocks downstream type-checked callers that already had widened parameter types and were forced into cast() or list(...) conversions.
pinecone/__init__.pyi stub for Jedi-based IDEs (DX-0145)Editors that use the Jedi static analyzer (Spyder, Eric, Cython, some older VS Code setups) don't traverse from … import * re-exports, so from pinecone import Pinecone showed no autocomplete in those environments. A new pinecone/__init__.pyi stub enumerates the public surface explicitly. mypy, Pyright, and modern Pylance users are unaffected (they already resolved through the runtime module); Jedi users now see full autocomplete and go-to-definition.
A CI typecheck-drift gate (DX-0146) ensures the stub stays in sync with the runtime pinecone/__init__.py.
PreviewTextQuery: field renamed to fieldsPreviewTextQuery.field is renamed to fields so it can accept multiple field names for multi-field text search. The v9.0.0 field= kwarg continues to work as a deprecated alias and will be removed in a future release. pc.preview is not covered by SemVer; see the v9.0.0 release notes.
These are not breaking under SemVer (they reject inputs the backend was already rejecting), but a few new client-side checks may surface as ValidationError where v9.0.0 surfaced as a server-side error. If your code wraps these calls in try/except, no changes are needed — ValidationError is a subclass of PineconeException — but the error class and message will differ.
top_k > 10000 on GrpcIndex.query is now rejected client-side. REST had this check already.top_n < 1 on pc.inference.rerank(...) is now rejected client-side.limit < 1 on Index.fetch_by_metadata is now rejected client-side (AGT-0044).upsert_records rejects records with an invalid _id field (non-string, empty, or containing null bytes) before the HTTP call.pc.collections.create(name=...) validates the resource name client-side (length, character set).Projects.create/update validates the project name (length ≤ 50, no null bytes) and refuses negative max_pods.ApiKeys.create enforces the 80-character name length cap.PreviewDocuments.fetch(ids=...) rejects empty ids lists.Two PreviewDocuments parameters that were never functional were removed:
PreviewDocuments.fetch(..., filter=...) — the filter parameter was dropped; it had no effect on the backend.PreviewDocuments.delete(..., filter=...) — same.If you were passing filter= to either of these on v9.0.0, the call will now raise TypeError. The argument was never being sent, so removing it does not change observable behavior.
fix(grpc): Cancelled + "Timeout expired" from the gRPC backend is now raised as PineconeTimeoutError instead of bubbling as the lower-level grpc.RpcError (CI-0052). Code that caught PineconeTimeoutError already handles this; code that caught only grpc.RpcError will no longer see those specific timeouts.
A number of v9.0.0 response structs had fields typed as non-optional that the backend in fact returns as null in some states, causing decode failures. These have been widened:
IndexModel.host — optional (null during creation / Terminating states)BackupModel.tags — dict[str, Any] (was dict[str, str])BackupModel.schema — real schema field (was an always-None stub)BackupModel.metric — removed (never returned by the backend)PodSpecInfo.replicas, .shards, .pods — optionalRestoreJobModel.created_at — optionalAPIKeyModel.name — str | None; the dead description field is removedStreamMessageEnd.usage — optional (handles null from streaming responses)PreviewBackupModel.tags — dict[str, Any] (was dict[str, str])PreviewIndexModel.host — optionalPreviewManagedDeployment.environment — optionalPreviewReadCapacityStatus.error_message — new field for failure statesModelInfoSupportedParameter.min / .max — int | float | None (v9.0.0 declared float, but the backend returns ints for some parameters)ByocSpecInfo.schema — newly surfacedIf you were destructuring these fields and bypassing None handling, add a guard. mypy will surface the new None possibilities under --strict.
schema and read_capacity through every index-creation path: integrated (pc.indexes.create_for_model), BYOC (build_byoc_body), ServerlessSpec, IntegratedSpec, the async creation path, and the create_index backcompat shim. On v9.0.0 these fields were dropped at the boundary depending on which path you called.pc.indexes.configure(...): stop pre-merging tags client-side; the backend handles tag merges, and the client-side step was overwriting partial updates. Same job, fewer bugs.pc.indexes.configure(...): relax validate_read_capacity to allow partial Dedicated patches (e.g. changing min_replicas without re-specifying max_replicas).configure boundary so that schema={"fields": {...}} and the shorter schema={"my_field": ...} are both accepted.IndexTerminatedError when an index enters Terminating or Disabled during a poll, instead of looping until the operator's timeout.EmbedConfig: include dimension in the create request (was silently dropped).pc.assistants.create(...): omit metadata from the request body when not provided. v9.0.0 sent metadata: null, which a stricter backend rejected.pc.assistants.list_page / list_files_page: upgrade to API version v202604, fix query-parameter names and pagination-token parsing, add page_size.metadata field as multipart form data (was being sent as a query param, which silently dropped the value).ChatCompletionMessage: role and content are now optional (matches the streaming assistant-role chunk shape where these arrive on a later chunk).ChatCompletionStreamChunk: expose the usage field, which had been dropped during struct decoding.EntailmentResult.reasoning: populate from the evaluate_alignment response (was always empty on v9.0.0).OperationModel.error: map from the JSON field error_message (was empty for failed operations).PINECONE_PLUGIN_ASSISTANT_CONTROL_HOST / PINECONE_PLUGIN_ASSISTANT_DATA_HOST env vars at client construction (legacy plugin compatibility).Terminated and Terminating as terminal states in _poll_until_ready, so creation polling exits promptly when an assistant transitions to a terminal failure state.Index.search / AsyncIndex.search: validate and normalize sparse/hybrid vector dict keys; wrap dense vector queries in {"values": ...} for backend serde compatibility.query_namespaces: preserve input namespace order in the result so ties break deterministically.upsert_records: drop the redundant id key when _id is present (the backend treats either, but sending both caused a 400).pc.backups.create(...): omit description from the request body when not provided.pc.indexes.create_from_backup(timeout=-1): return CreateIndexFromBackupResponse directly instead of attempting to poll with a negative timeout.limit=10 default on Backups.list / RestoreJobs.list so backend-default pagination is used.PreviewIndexes.create and configure: full tag validation (was incomplete on v9.0.0).PreviewIndexes.configure: validate schema fields are semantic_text only.PreviewIndexes.list_backups: forward limit to the server in fetch_page (was hardcoded).PreviewBooleanField and PreviewLegacyIntegerField are now in the schema union; v9.0.0 rejected schemas containing them.Projects.create / update: client-side name length and null-byte validation.Projects.update: reject negative max_pods before the HTTP call.ApiKeys.create: enforce the 80-character name length cap.ApiKeys.create: remove the dead description parameter (server doesn't accept it).rerank(...): validate top_n >= 1 before the HTTP call.list_models(...): reject type=rerank combined with a vector_type argument.ModelInfoSupportedParameter.min/max typing widened to int | float | None.import_data(...): omit errorMode from the wire payload when error_mode is not specified (was sending null, which the backend rejected).schema and read_capacity through async index-creation shims so that v8-style positional calls continue to behave correctly.AssistantModel.context shim to accept a messages= argument (matches the deprecated v8 signature).A broad pass over docstrings, RST formatting, and the v9 migration guide. Highlights:
pc.index() rather than the deprecated shim:meta private: so they no longer clutter generated docsPinecone.index, Pinecone.config, Pinecone.close, Pinecone.assistant, Pinecone.backups, Pinecone.collections, Pinecone.assistants: complete Args, Returns, Raises, and Examples sections where missingupsert_from_dataframe max_concurrency claim corrected in performance guidegrpc extra_AssistantNamespaceProxy; build remains warnings-as-errors cleanquery=, spec_embed=, field=), unsafe .get() calls, deprecated import paths, false immutability claimsThe Sphinx docs build remains green with -W (warnings-as-errors).
pytest-xdist (-n 6 --dist=loadfile); the release-readiness job's timeout was raised from 60 → 90 min as a margin.abi3audit pass runs on every wheel produced by the cross-platform matrix; catches stable-ABI violations before publish (the multi-Python smoke job is the runtime complement).abi3audit and auditwheel in their own venvs so they no longer leak into the verify Python's pip check.pinecone/ tree before running so the installed wheel is actually exercised.dev-version resolution now uses the artifacts API directly rather than walking workflow runs (avoids missing the most recent real publish when intermediate runs are no-op skips).If something worked on v9.0.0 and doesn't on v9.0.1, please open an issue at https://github.com/pinecone-io/python-sdk/issues with the same details requested in the v9.0.0 notes (versions, minimal repro, expected vs actual). Compatibility regressions are prioritized.
Full Changelog: v9.0.0...v9.0.1
Nothing published for this version
Most v8 code paths continue to work. Where signatures changed, deprecated aliases are in place. The migration guide enumerates the cases that need cod…
v9 is a total rewrite of the Pinecone Python SDK. Rewrites are always ambitious undertakings, and we were motivated by three outcomes that had become difficult to achieve incrementally on the v8 codebase:
pip install covering all transports, with a much smaller dependency tree.pc.preview namespace introduced in this release is a concrete example — it would not have been feasible for us to ship in the v8 client. This benefit is harder to quantify than installation or latency, but it changes the cost of every future feature, which adds up over time.We made an effort to preserve much of the public surface of the SDK. Most v8 code is expected to continue to run unchanged — but the internals are entirely new. If you are upgrading from v8, start with the migration guide.
pip install pinecone. The [grpc] and [asyncio] extras are no longer needed; both transports ship in the base package.mypy --strict is clean; IDE autocomplete and downstream type-checked codebases see complete annotations.httpx, msgspec, orjson. v8.1.2 declared 7 in its base install, plus up to 7 more across the [grpc] and [asyncio] extras (14 in a fully-enabled install). A smaller dependency tree means fewer version conflicts in your environment and a smaller third-party security-advisory surface to track.pc.indexes, pc.collections, pc.backups, pc.inference, pc.assistant, and pc.preview mirror the resources they act on. The flat v8 method names are preserved as aliases.pinecone-plugin-assistant package and the plugin discovery system are retired; pc.assistant is part of the core client.pc.preview introduces a namespace for public preview features, beginning with full-text search over documents.pip install pinecone
This is the entire install for sync REST, asyncio REST, and gRPC. The gRPC transport is now a Rust extension built into the wheel, so there is no grpcio to install, no version pinning or conflicts to manage with other dependencies of your app.
from pinecone import Pinecone
pc = Pinecone(api_key="...")
index = pc.index("my-index") # sync REST
grpc_index = pc.index("my-index", grpc=True) # gRPC, no grpcio dependency
Python 3.10+ is required. See Migrating from v8 for why Python 3.9 was dropped.
The improvements come from three changes to the internals:
msgspec and orjson. Response objects are typed structs decoded at native speed, not dicts populated by Python-level loops.The numbers below are from end-to-end benchmarks against a real Pinecone serverless index (1536-dim, AWS us-east-1, p50 latencies). Ratios are v8 / v9. Your mileage will vary with index configuration, region, client hardware, and network—treat these figures as directional, not a guarantee.
| Scenario | v8 | v9 | Ratio |
|---|---|---|---|
| Upsert b=100 (sync REST) | 1.00 s | 306 ms | 3.3× |
| Upsert b=100 (async REST) | 1.01 s | 308 ms | 3.3× |
| Throughput 10k vectors @ concurrency=20 (async) | 64.6 s | 4.45 s | 14.5× |
| Throughput 10k vectors @ concurrency=100 (sync) | 75.1 s | 4.42 s | 17.0× |
| Query top_k=100, +values +metadata | 786 ms | 287 ms | 2.7× |
| Query top_k=1000, +values | 7.01 s | 2.05 s | 3.4× |
Cold import (python -c "import pinecone") |
196 ms | 17.6 ms | 11.2× |
| Net cold start (import + construct) | ~210 ms | ~45 ms | ~5× |
The largest practical change is in bulk upsert under concurrency. v8's REST path saturates client CPU on serialization and stops scaling past a single connection — adding workers does not increase throughput. v9 sustains roughly 16× the v8 sync throughput at concurrency=20 and lands within 2× of gRPC. If you adopted gRPC primarily for upsert throughput on REST, the REST path in v9 may be fast enough to remove that complexity from your stack.
gRPC per-call latency is roughly at parity with v8's Python grpcio channel, which is expected — the wire format dominates. The Rust channel scales further under high concurrency without GIL contention.
You can read a little more about this on Batching Large Upserts
mypy --strict runs clean across the codebase. The public surface — every class, method, parameter, and return value — is fully annotated. A handful of Any types remain at the JSON boundary, where they are the appropriate choice; the rest of the surface gives complete IDE autocomplete and works cleanly in downstream type-checked codebases.
v9 ships with three runtime dependencies: httpx[http2], msgspec, and orjson.
For comparison, v8.1.2 declared:
[grpc] extra[asyncio] extra— a maximum of 14 declared runtime dependencies for a fully-enabled install. v9 includes both gRPC and async support in the base package and declares only three. The practical benefits:
v9’s control plane is resource-oriented: methods are grouped under the resource they target—pc.indexes, pc.collections, pc.backups, pc.inference, and so on—instead of a long flat list on the root client. That lines up with the API’s resource model and will make it easier to maintain and navigate as we continue adding new capabilities.
pc.indexes.create(name="my-index", dimension=1536, spec=...)
pc.indexes.list()
pc.collections.create(...)
pc.backups.list()
pc.inference.embed(...)
This is mostly an organizational change, but it matters as the surface grows. Piling every method onto the root client scales poorly; giving each resource its own subtree keeps discovery sane and gives new areas—pc.assistant, pc.preview, and whatever comes next—a stable place to land without fighting for method names on pc.
Existing v8 code does not need to change. The flat v8 method names are preserved as aliases on the client:
# Both forms work in v9
pc.create_index(name="my-index", dimension=1536, spec=...)
pc.indexes.create(name="my-index", dimension=1536, spec=...)
pc.list_indexes()
pc.indexes.list()
The resource-oriented (namespaced) form is recommended for new code.
The Pinecone Assistant API previously shipped as a separate plugin (pinecone-plugin-assistant) installed alongside pinecone. In v9, it is part of the main package. Code that imported from pinecone_plugins.assistant.* should switch to pinecone.models.assistant:
from pinecone import Pinecone
from pinecone.models.assistant import Message
pc = Pinecone(api_key="...")
assistant = pc.assistant.create_assistant("my-assistant")
assistant.upload_file(file_path="report.pdf")
response = assistant.chat(messages=[Message(role="user", content="...")])
The runtime methods on pc.assistant (create_assistant, list_assistants, describe_assistant, chat, upload_file, …) are unchanged. Only the import paths moved.
The plugin discovery system itself is retired. Going forward, Assistant improvements ship in the main SDK release stream rather than as separately-versioned packages, which removes a coordination step and shortens the path from feature work to general availability.
The full v8-to-v9 import-path mapping is in §8 of the migration guide.
pc.preview: a namespace for Early Access and Public Preview featuresv9 introduces pc.preview, a dedicated namespace for features that are exposed in the SDK but still in active development. The first occupant is document upsert and full-text search:
from pinecone import Pinecone
pc = Pinecone(api_key="...")
index = pc.preview.index(name="articles-en-preview")
results = index.documents.search(
namespace="articles-en",
top_k=5,
score_by=[{"field": "embedding", "query": [0.012, -0.087, 0.153]}],
)
[!NOTE] Anything reachable through
pc.previewis not covered by SemVer. Signatures, behavior, and availability may change in any minor SDK release, and the SDK's normal deprecation policy does not apply. Pin the SDK version if your code depends on a preview feature.
Most v8 code paths run on v9 unchanged, and where signatures changed, deprecated aliases keep the old names working. The cases below are the ones most likely to require code changes:
PineconeAsyncio is renamed to AsyncPinecone. The old name still imports as a deprecated alias.msgspec.Struct instances, not dicts. Field access (idx.name, idx.dimension) is unchanged, but dict(idx) raises. Use msgspec.structs.asdict(idx) if you need a dict.batch_size= is set, failures are captured on the response (response.has_errors, response.failed_items) rather than raised. Code that wraps batched upserts in try/except will silently undercount unless updated. Single-request upsert (the default) keeps the v8 raise-on-failure behavior.Pinecone(retries=3) becomes Pinecone(retry_config=RetryConfig(max_retries=3)).pinecone.core.client.api… no longer resolve. Use top-level pinecone imports.pinecone_plugins.assistant.* imports are removed. See the Assistant section above.The full mapping table and code-level migration notes are in the migration guide.
v9 went through an extensive unit and integration test suite, end-to-end exercises against production Pinecone services, and a benchmark suite that re-runs the v8 surface against v9 to catch behavioral regressions. Real-world usage always exercises code paths that internal testing doesn't, though, and reports from early upgraders are the fastest path to a polished v9.
If something worked on v8.1.2 and doesn't on v9, please open an issue at https://github.com/pinecone-io/python-sdk/issues with:
Compatibility regressions are prioritized, and the smaller the repro, the faster we can ship a fix. Performance observations, install challenges, type-checking issues, migration-guide gaps, namespace-pattern rough edges — anything that surprises you — are welcome on the same tracker.
Full Changelog: https://github.com/pinecone-io/python-sdk/compare/v8.1.2...v9.0.0
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fix AttributeError on delete with non-dict response (#564) by @jhamon in https://github.com/pinecone-io/pinecone-python-client/pull/634
Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v8.1.1...v8.1.2
Full Changelog: v8.1.1...v8.1.2
Bump orjson minimum to 3.11.6 (CVE-2025-67221)
Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v8.1.0...v8.1.1
This release adds support for creating and configuring index read_capacity for BYOC indexes:
This release adds support for creating and configuring index read_capacity for BYOC indexes:
import pinecone
from pinecone import ByocSpec
pc = pinecone.Pinecone(api_key="YOUR_API_KEY")
# Create a BYOC index with OnDemand read capacity
pc.create_index(
name="my-byoc-index",
dimension=1536,
spec=ByocSpec(
environment="my-byoc-env",
read_capacity={"mode": "OnDemand"},
)
)
# Create a BYOC index with Dedicated read capacity
pc.create_index(
name="my-byoc-index",
dimension=1536,
spec=ByocSpec(
environment="my-byoc-env",
read_capacity={
"mode": "Dedicated",
"dedicated": {
"node_type": "b1",
"scaling": "Manual",
"manual": {"replicas": 2},
},
},
)
)
The following user-facing types have been added or updated to support this:
ByocSpec — now accepts optional read_capacity and schema fieldsReadCapacityDict — union alias for the two read capacity modes belowReadCapacityOnDemandDict — {"mode": "OnDemand"}ReadCapacityDedicatedDict — {"mode": "Dedicated", "dedicated": ReadCapacityDedicatedConfigDict}ReadCapacityDedicatedConfigDict — {"node_type": str, "scaling": str, "manual": ScalingConfigManualDict}ScalingConfigManualDict — {"shards": int, "replicas": int}MetadataSchemaFieldConfig — {"filterable": bool}, used with the schema field on ByocSpecAll of the above are exported from the top-level pinecone module.
Support for scan_factor and max_candidates has been added to Index.query() and Index.query_namespaces():
# scan_factor widens the IVF scan to trade latency for higher recall
# max_candidates controls how many candidates are reranked with exact distances
results = index.query(
vector=[...],
top_k=10,
scan_factor=2.0,
max_candidates=500,
)
Both parameters are optional and only take effect on dedicated read node (DRN) dense indexes. scan_factor adjusts how much of the IVF index is scanned when gathering vector candidates, and max_candidates caps the number of candidates that undergo exact-distance reranking to improve recall.
2025-10, implement schema/read_capacity in BYOCSpec by @austin-denoble in https://github.com/pinecone-io/pinecone-python-client/pull/614scan_factor and max_candidates for query by @austin-denoble in https://github.com/pinecone-io/pinecone-python-client/pull/617Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v8.0.1...v8.1.0
Nothing published for this version
🔒 Fixed Protobuf Denial-of-Service Vulnerability (CVE-2025-4565)
🔒 Fixed Protobuf Denial-of-Service Vulnerability (CVE-2025-4565)
Updated protobuf dependency to address a denial-of-service vulnerability when parsing deeply nested recursive structures in a Pure-Python backend.
Affected users: Only users of the grpc extras (pip install pinecone[grpc]) and PineconeGRPC client will be affected by the change. Users of the default REST client (Pinecone) are not affected.
Changes:
protobuf from 5.x to 6.33.0+googleapis-common-protos from 1.66.0 to 1.72.0+ for compatibilityImpact:
6.33.0 (was 5.29.5)<6.33.0 will need to upgradeReferences:
Dependency updates: Updated protobuf to 5.29.5 to address security vulnerabilities.
7.x to 8.xThe v8 release of the Pinecone Python SDK has been published as pinecone to PyPI.
With a few exceptions noted below, nearly all changes are additive and non-breaking. The major version bump primarily reflects the step up to API version 2025-10 and the addition of a new dependency on orjson for fast JSON parsing.
⚠️ Python 3.9 is no longer supported. The SDK now requires Python 3.10 or later. Python 3.9 reached end-of-life on October 2, 2025. Users must upgrade to Python 3.10+ to continue using the SDK.
⚠️ Namespace parameter default behavior changed. The SDK no longer applies default values for the namespace parameter in GRPC methods. When namespace=None, the parameter is omitted from requests, allowing the API to handle namespace defaults appropriately. This change affects upsert_from_dataframe methods in GRPC clients. The API is moving toward "__default__" as the default namespace value, and this change ensures the SDK doesn't override API defaults.
Note: The official SDK package was renamed last year from pinecone-client to pinecone beginning in version 5.1.0. Please remove pinecone-client from your project dependencies and add pinecone instead to get the latest updates if upgrading from earlier versions.
8.xYou can now configure dedicated read nodes for your serverless indexes, giving you more control over query performance and capacity planning. By default, serverless indexes use OnDemand read capacity, which automatically scales based on demand. With dedicated read capacity, you can allocate specific read nodes with manual scaling control.
Create an index with dedicated read capacity:
from pinecone import (
Pinecone,
ServerlessSpec,
CloudProvider,
AwsRegion,
Metric
)
pc = Pinecone()
pc.create_index(
name='my-index',
dimension=1536,
metric=Metric.COSINE,
spec=ServerlessSpec(
cloud=CloudProvider.AWS,
region=AwsRegion.US_EAST_1,
read_capacity={
"mode": "Dedicated",
"dedicated": {
"node_type": "t1",
"scaling": "Manual",
"manual": {
"shards": 2,
"replicas": 2
}
}
}
)
)
Configure read capacity on an existing index:
You can switch between OnDemand and Dedicated modes, or adjust the number of shards and replicas for dedicated read capacity:
from pinecone import Pinecone
pc = Pinecone()
# Switch to OnDemand read capacity
pc.configure_index(
name='my-index',
read_capacity={"mode": "OnDemand"}
)
# Switch to Dedicated read capacity with manual scaling
pc.configure_index(
name='my-index',
read_capacity={
"mode": "Dedicated",
"dedicated": {
"node_type": "t1",
"scaling": "Manual",
"manual": {
"shards": 3,
"replicas": 2
}
}
}
)
# Scale up by increasing shards and replicas
pc.configure_index(
name='my-index',
read_capacity={
"mode": "Dedicated",
"dedicated": {
"node_type": "t1",
"scaling": "Manual",
"manual": {
"shards": 4,
"replicas": 3
}
}
}
)
When you change read capacity configuration, the index will transition to the new configuration. You can use describe_index to check the status of the transition.
See PR #528 for details.
You can now fetch vectors using metadata filters instead of vector IDs. This is especially useful when you need to retrieve vectors based on their metadata properties.
from pinecone import Pinecone
pc = Pinecone()
index = pc.Index(host="your-index-host")
# Fetch vectors matching a complex filter
response = index.fetch_by_metadata(
filter={'genre': {'$in': ['comedy', 'drama']}, 'year': {'$eq': 2019}},
namespace='my_namespace',
limit=50
)
print(f"Found {len(response.vectors)} vectors")
# Iterate through fetched vectors
for vec_id, vector in response.vectors.items():
print(f"ID: {vec_id}, Metadata: {vector.metadata}")
Pagination support:
When fetching large numbers of vectors, you can use pagination tokens to retrieve results in batches:
# First page
response = index.fetch_by_metadata(
filter={'status': 'active'},
limit=100
)
# Continue with next page if available
if response.pagination and response.pagination.next:
next_response = index.fetch_by_metadata(
filter={'status': 'active'},
pagination_token=response.pagination.next,
limit=100
)
The update method used to require a vector id to be passed, but now you have the option to pass a metadata filter instead. This is useful for bulk metadata updates across many vectors.
There is also a dry_run option that allows you to preview the number of vectors that would be changed by the update before performing the operation.
from pinecone import Pinecone
pc = Pinecone()
index = pc.Index(host="your-index-host")
# Preview how many vectors would be updated (dry run)
response = index.update(
set_metadata={'status': 'active'},
filter={'genre': {'$eq': 'drama'}},
dry_run=True
)
print(f"Would update {response.matched_records} vectors")
# Apply the update by repeating the command without dry_run
response = index.update(
set_metadata={'status': 'active'},
filter={'genre': {'$eq': 'drama'}}
)
A new FilterBuilder utility class provides a type-safe, fluent interface for constructing metadata filters. While perhaps a bit verbose, it can help prevent common errors like misspelled operator names and provides better IDE support.
When you chain .build() onto the FilterBuilder it will emit a python dictionary representing the filter. Methods that take metadata filters as arguments will continue to accept dictionaries as before.
from pinecone import Pinecone, FilterBuilder
pc = Pinecone()
index = pc.Index(host="your-index-host")
# Simple equality filter
filter1 = FilterBuilder().eq("genre", "drama").build()
# Returns: {"genre": "drama"}
# Multiple conditions with AND using & operator
filter2 = (FilterBuilder().eq("genre", "drama") &
FilterBuilder().gt("year", 2020)).build()
# Returns: {"$and": [{"genre": "drama"}, {"year": {"$gt": 2020}}]}
# Multiple conditions with OR using | operator
filter3 = (FilterBuilder().eq("genre", "comedy") |
FilterBuilder().eq("genre", "drama")).build()
# Returns: {"$or": [{"genre": "comedy"}, {"genre": "drama"}]}
# Complex nested conditions
filter4 = ((FilterBuilder().eq("genre", "drama") &
FilterBuilder().gte("year", 2020)) |
(FilterBuilder().eq("genre", "comedy") &
FilterBuilder().lt("year", 2000))).build()
# Use with fetch_by_metadata
response = index.fetch_by_metadata(filter=filter2, limit=50)
# Use with update
index.update(
set_metadata={'status': 'archived'},
filter=filter3
)
The FilterBuilder supports all Pinecone filter operators: eq, ne, gt, gte, lt, lte, in_, nin, and exists. Compound expressions are built with and as & and or as |.
See PR #529 for fetch_by_metadata, PR #544 for update() with filter, and PR #531 for FilterBuilder.
You can now create namespaces in serverless indexes directly from the SDK:
from pinecone import Pinecone
pc = Pinecone()
index = pc.Index(host="your-index-host")
# Create a namespace with just a name
namespace = index.create_namespace(name="my-namespace")
print(f"Created namespace: {namespace.name}, Vector count: {namespace.vector_count}")
# Create a namespace with schema configuration
namespace = index.create_namespace(
name="my-namespace",
schema={
"fields": {
"genre": {"filterable": True},
"year": {"filterable": True}
}
}
)
Note: This operation is not supported for pod-based indexes.
See PR #532 for details.
For sparse indexes with integrated embedding configured to use the pinecone-sparse-english-v0 model, you can now specify which terms must be present in search results:
from pinecone import Pinecone, SearchQuery
pc = Pinecone()
index = pc.Index(host="your-index-host")
response = index.search(
namespace="my-namespace",
query=SearchQuery(
inputs={"text": "Apple corporation"},
top_k=10,
match_terms={
"strategy": "all",
"terms": ["apple", "corporation"]
}
)
)
The match_terms parameter ensures that all specified terms must be present in the text of each search hit. Terms are normalized and tokenized before matching, and order does not matter.
See PR #530 for details.
Update API keys, projects, and organizations:
from pinecone import Admin
admin = Admin() # Auth with PINECONE_CLIENT_ID and PINECONE_CLIENT_SECRET
# Update an API key's name and roles
api_key = admin.api_key.update(
api_key_id='my-api-key-id',
name='updated-api-key-name',
roles=['ProjectEditor', 'DataPlaneEditor']
)
# Update a project's configuration
project = admin.project.update(
project_id='my-project-id',
name='updated-project-name',
max_pods=10,
force_encryption_with_cmek=True
)
# Update an organization
organization = admin.organization.update(
organization_id='my-org-id',
name='updated-organization-name'
)
Delete organizations:
from pinecone import Admin
admin = Admin()
# Delete an organization (use with caution!)
admin.organization.delete(organization_id='my-org-id')
See PR #527 and PR #543 for details.
You can now configure which metadata fields are filterable when creating serverless indexes. This helps optimize performance by only indexing metadata fields that you plan to use for filtering:
from pinecone import (
Pinecone,
ServerlessSpec,
CloudProvider,
AwsRegion,
Metric
)
pc = Pinecone()
pc.create_index(
name='my-index',
dimension=1536,
metric=Metric.COSINE,
spec=ServerlessSpec(
cloud=CloudProvider.AWS,
region=AwsRegion.US_EAST_1,
schema={
"genre": {"filterable": True},
"year": {"filterable": True},
"rating": {"filterable": True}
}
)
)
When using schemas, only fields marked as filterable: True in the schema can be used in metadata filters.
See PR #528 for details.
The SDK now exposes header information from API responses. This information is available in response objects via the _response_info attribute and can be useful for debugging and monitoring.
from pinecone import Pinecone
pc = Pinecone()
index = pc.Index(host="your-index-host")
# Perform a query
response = index.query(
vector=[0.1, 0.2, 0.3, ...],
top_k=10,
namespace='my_namespace'
)
# Access response headers
if hasattr(response, '_response_info') and response._response_info:
headers = response._response_info.get('raw_headers', {})
# Access specific headers (header names are normalized to lowercase)
lsn = headers.get('x-pinecone-request-lsn')
if lsn:
print(f"LSN: {lsn}")
# View all available headers
for header_name, header_value in headers.items():
print(f"{header_name}: {header_value}")
See PR #539 for details.
We've replaced Python's standard library json module with orjson, a fast JSON library written in Rust. This provides significant performance improvements for both serialization and deserialization of request payloads:
These improvements are especially beneficial for:
No code changes are required - the API remains the same, and you'll automatically benefit from these performance improvements.
See PR #556 for details.
We've optimized gRPC response parsing by replacing json_format.MessageToDict with direct protobuf field access. This optimization provides approximately 2x faster response parsing for gRPC operations.
Special thanks to @yorickvP for surfacing the json_format.MessageToDict refactor opportunity. While we didn't merge the specific PR, yorick's insight led us to implement a similar optimization that significantly improves gRPC performance.
See PR #553 for details.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
This minor release includes the ability to interact with the Admin API and adds support for working with index namespaces via gRPC. Previously, namesp
This minor release includes the ability to interact with the Admin API and adds support for working with index namespaces via gRPC. Previously, namespace support was available only through REST.
This release introduces an Admin class that provides support for performing CRUD operations on projects and API keys using REST.
from pinecone import Admin
# Use service account credentials
admin = Admin(client_id='foo', client_secret='bar')
# Example: Create a project
project = admin.project.create(
name="example-project",
max_pods=5
)
print(f"Project {project.id} was created")
# Example: Rename a project
project = admin.project.get(name='example-project')
admin.project.update(
project_id=project.id,
name='my-awesome-project'
)
# Example: Enable CMEK on all projects
project_list = admin.projects.list()
for proj in project_list_response.data:
admin.projects.update(
project_id=proj.id,
force_encryption_with_cmek=True
)
# Example: Set pod quota to 0 for all projects
project_list = admin.projects.list()
for proj in project_list_response.data:
admin.projects.update(project_id=proj.id, max_pods=0)
# Delete the project
admin.project.delete(project_id=project.id)
from pinecone import Admin
# Use service account credentials
admin = Admin(client_id='foo', client_secret='bar')
project = admin.project.get(name='my-project')
# Create an API key
api_key_response = admin.api_keys.create(
project_id=project.id,
name="ci-key",
roles=["ProjectEditor"]
)
key = api_key_response.value # 'pcsk_....'
# Look up info on a key by id
key_info = admin.api_keys.get(
api_key_id=api_key_response.key.id
)
# Delete a key
admin.api_keys.delete(
api_key_id=api_key_response.key.id
)
The gRPC Index class now exposes methods for calling describe_namespace, delete_namespace, list_namespaces, and list_namespaces_paginated.
from pinecone.grpc import PineconeGRPC as Pinecone
pc = Pinecone(api_key='YOUR_API_KEY')
index = pc.Index(host='your-index-host')
# list namespaces
results = index.list_namespaces_paginated(limit=10)
next_results = index.list_namespaces_paginated(limit=10, pagination_token=results.pagination.next)
# describe namespace
namespace = index.describe_namespace(results[0].name)
# delete namespaces (NOTE: this deletes all data within the namespace)
index.delete_namespace(results[0].name)
Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v7.2.0...v7.3.0
This minor release includes new methods for working with index namespaces via REST, and the ability to configure an index with the embed configuration
This minor release includes new methods for working with index namespaces via REST, and the ability to configure an index with the embed configuration, which was not previously exposed.
The Index and IndexAsyncio classes now expose methods for calling describe_namespace, delete_namespace, list_namespaces, and list_namespaces_paginated. There is also a NamespaceResource which can be used to perform these operations. Namespaces themselves are still created implicitly when upserting data to a specific namespace.
from pinecone import Pinecone
pc = Pinecone(api_key='YOUR_API_KEY')
index = pc.Index(host='your-index-host')
# list namespaces
results = index.list_namespaces_paginated(limit=10)
next_results = index.list_namespaces_paginated(limit=10, pagination_token=results.pagination.next)
# describe namespace
namespace = index.describe_namespace(results[0].name)
# delete namespaces (NOTE: this deletes all data within the namespace)
index.delete_namespace(results[0].name)
Previously, the configure_index methods did not support providing an embed argument when configuring an existing index. These methods now support embed in the shape of ConfigureIndexEmbed. You can convert an existing index to an integrated index by specifying the embedding model and field_map. The index vector type and dimension must match the model vector type and dimension, and the index similarity metric must be supported by the model. You can use list_models and get_model on the Inference class to get specific details about models.
You can later change the embedding configuration to update the field map, read parameters, or write parameters. Once set, the model cannot be changed.
from pinecone import Pinecone
pc = Pinecone(api_key='YOUR_API_KEY')
# convert an existing index to use the integrated embedding model multilingual-e5-large
pc.configure_index(
name="my-existing-index",
embed={"model": "multilingual-e5-large", "field_map": {"text": "chunk_text"}},
)
embed to Index configure calls by @austin-denoble in https://github.com/pinecone-io/pinecone-python-client/pull/515Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v7.1.0...v7.2.0
This release fixes an issue where GRPC methods using async_req=True ignored user-provided timeout values, defaulting instead to a hardcoded 5-second t
This release fixes an issue where GRPC methods using async_req=True ignored user-provided timeout values, defaulting instead to a hardcoded 5-second timeout imposed by PineconeGrpcFuture. To verify this fix, we added a new test file, test_timeouts.py, which uses a mock GRPC server to simulate client timeout behavior under delayed response conditions.
Nothing published for this version
This small bugfix release includes the following fixes:
This small bugfix release includes the following fixes:
readline error reported by in #502. See #503 for details on the root cause and fix.pinecone-plugin-assistant. The assistant plugin had been inadvertently added as a dev dependency rather than a dependency, which means our integration tests for that functionality were able to pass while the published artifact was not including it. We have corrected this problem, which means assistant functions should now work without installing anything additional.Nothing published for this version
This small bugfix release fixes:
This small bugfix release fixes:
Nothing published for this version
There are no intentional breaking changes between v6 and v7 of the SDK. The major version bump reflects the move from calling the 2025-01 to the 2025-…
6.x to 7.xThe v7 release of the Pinecone Python SDK has been published as pinecone to PyPI.
There are no intentional breaking changes between v6 and v7 of the SDK. The major version bump reflects the move from calling the 2025-01 to the 2025-04 version of the underlying API.
Some internals of the client have been reorganized or moved, but we've made an effort to alias everything and show warning messages when appropriate. If you experience any unexpected breaking changes that cause you friction while upgrading, let us know and we'll try to smooth it out.
[!NOTE] The official SDK package was renamed from
pinecone-clienttopineconebeginning in version 5.1.0. Please removepinecone-clientfrom your project dependencies and addpineconeinstead to get the latest updates if upgrading from earlier versions.
7.xNew Features:
Other useful improvements:
py.typed marker file to indicate inline type information is present in the package. We're still working toward reaching full coverage with our type hints, but including this file allows some tools to find the inline definitions we have already implemented.You can create backups from your Serverless indexes and use these backups to create new indexes. Some fields such as record_count are initially empty but will be populated by the time a backup is ready for use.
from pinecone import Pinecone
pc = Pinecone()
index_name = 'example-index'
if not pc.has_index(name=index_name):
raise Exception('An index must exist before backing it up')
backup = pc.create_backup(
index_name=index_name,
backup_name='example-backup',
description='testing out backups'
)
# {
# "backup_id": "4698a618-7e56-4a44-93bc-fc8f1371aa36",
# "source_index_name": "example-index",
# "source_index_id": "ec6fd44c-ab45-4873-97f3-f6b44b67e9bc",
# "status": "Initializing",
# "cloud": "aws",
# "region": "us-east-1",
# "tags": {},
# "name": "example-backup",
# "description": "testing out backups",
# "dimension": null,
# "record_count": null,
# "namespace_count": null,
# "size_bytes": null,
# "created_at": "2025-05-16T18:44:28.480671533Z"
# }
Check the status of a backup
from pinecone import Pinecone
pc = Pinecone()
pc.describe_backup(backup_id='4698a618-7e56-4a44-93bc-fc8f1371aa36')
# {
# "backup_id": "4698a618-7e56-4a44-93bc-fc8f1371aa36",
# "source_index_name": "example-index",
# "source_index_id": "ec6fd44c-ab45-4873-97f3-f6b44b67e9bc",
# "status": "Ready",
# "cloud": "aws",
# "region": "us-east-1",
# "tags": {},
# "name": "example-backup",
# "description": "testing out backups",
# "dimension": 768,
# "record_count": 1000,
# "namespace_count": 1,
# "size_bytes": 289656,
# "created_at": "2025-05-16T18:44:28.480691Z"
# }
You can use list_backups to see all of your backups and their current status. If you have a large number of backups, results will be paginated. You can control the pagination with optional parameters for limit and pagination_token.
from pinecone import Pinecone
pc = Pinecone()
# All backups
pc.list_backups()
# Only backups associated with a particular index
pc.list_backups(index_name='my-index')
To create an index from a backup, use create_index_from_backup.
from pinecone import Pinecone
pc = Pinecone()
pc.create_index_from_backup(
name='index-from-backup',
backup_id='4698a618-7e56-4a44-93bc-fc8f1371aa36',
deletion_protection = "disabled",
tags={'env': 'testing'},
)
Under the hood, a restore job is created to handle taking data from your backup and loading it into the newly created serverless index. You can check status of pending restore jobs with pc.list_restore_jobs() or pc.describe_restore_job()
You can now fetch a dynamic list of models supported by the Inference API.
from pinecone import Pinecone
pc = Pinecone()
# List all models
models = pc.inference.list_models()
# List models, with model type filtering
models = pc.inference.list_models(type="embed")
models = pc.inference.list_models(type="rerank")
# List models, with vector type filtering
models = pc.inference.list_models(vector_type="dense")
models = pc.inference.list_models(vector_type="sparse")
# List models, with both type and vector type filtering
models = pc.inference.list_models(type="rerank", vector_type="dense")
Or, if you know the name of a model, you can get just those details
pc.inference.get_model(model_name='pinecone-rerank-v0')
# {
# "model": "pinecone-rerank-v0",
# "short_description": "A state of the art reranking model that out-performs competitors on widely accepted benchmarks. It can handle chunks up to 512 tokens (1-2 paragraphs)",
# "type": "rerank",
# "supported_parameters": [
# {
# "parameter": "truncate",
# "type": "one_of",
# "value_type": "string",
# "required": false,
# "default": "END",
# "allowed_values": [
# "END",
# "NONE"
# ]
# }
# ],
# "modality": "text",
# "max_sequence_length": 512,
# "max_batch_size": 100,
# "provider_name": "Pinecone",
# "supported_metrics": []
# }
For customers using our BYOC offering, you can now create indexes and list/describe indexes you have created in your cloud.
from pinecone import Pinecone, ByocSpec
pc = Pinecone()
pc.create_index(
name='example-byoc-index',
dimension=768,
metric='cosine',
spec=ByocSpec(environment='my-private-env'),
tags={
'env': 'testing'
},
deletion_protection='enabled'
)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
[Fix] Error when fetching sparse vector by id over grpc by @jhamon in https://github.com/pinecone-io/pinecone-python-client/pull/467
Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v6.0.1...v6.0.2
Nothing published for this version
This release contains a small fix to correct an incompatibility between the 6.0.0 pinecone release and pinecone-plugin-assistant. While working toward
This release contains a small fix to correct an incompatibility between the 6.0.0 pinecone release and pinecone-plugin-assistant. While working toward improving type coverage of the sdk, some attributes of the internal Configuraiton class were erroneously removed even though they are still needed by the plugin to load correctly.
The 6.0.0 pinecone SDK should now work with all versions of pinecone-plugin-assistant except for 1.1.1 which errors when used with 6.0.0.
Thanks @avi1mizrahi for contributing the fix.
Nothing published for this version
Removed some previously deprecated and rarely used keyword arguments (config, openapi_config, and index_api) to instead prefer dedicated keyword argum…
This release adds a new create_index_for_model method as well as upsert_records, and search methods. Together these methods provide a way for you to easily store your data and let us manage the process of creating embeddings. To learn about available models, see the Model Gallery.
Note: If you were previously using the preview versions of this functionality via the pinecone-plugin-records package, you will need to uninstall that package in order to use the v6 pinecone release.
from pinecone import (
Pinecone,
CloudProvider,
AwsRegion,
EmbedModel,
)
# 1. Instantiate the Pinecone client
pc = Pinecone(api_key="<<PINECONE_API_KEY>>")
# 2. Create an index configured for use with a particular model
index_config = pc.create_index_for_model(
name="my-model-index",
cloud=CloudProvider.AWS,
region=AwsRegion.US_EAST_1,
embed=IndexEmbed(
model=EmbedModel.Multilingual_E5_Large,
field_map={"text": "my_text_field"}
)
)
# 3. Instantiate an Index client
idx = pc.Index(host=index_config.host)
# 4. Upsert records
idx.upsert_records(
namespace="my-namespace",
records=[
{
"_id": "test1",
"my_text_field": "Apple is a popular fruit known for its sweetness and crisp texture.",
},
{
"_id": "test2",
"my_text_field": "The tech company Apple is known for its innovative products like the iPhone.",
},
{
"_id": "test3",
"my_text_field": "Many people enjoy eating apples as a healthy snack.",
},
{
"_id": "test4",
"my_text_field": "Apple Inc. has revolutionized the tech industry with its sleek designs and user-friendly interfaces.",
},
{
"_id": "test5",
"my_text_field": "An apple a day keeps the doctor away, as the saying goes.",
},
{
"_id": "test6",
"my_text_field": "Apple Computer Company was founded on April 1, 1976, by Steve Jobs, Steve Wozniak, and Ronald Wayne as a partnership.",
},
],
)
# 5. Search for similar records
from pinecone import SearchQuery, SearchRerank, RerankModel
response = index.search_records(
namespace="my-namespace",
query=SearchQuery(
inputs={
"text": "Apple corporation",
},
top_k=3
),
rerank=SearchRerank(
model=RerankModel.Bge_Reranker_V2_M3,
rank_fields=["my_text_field"],
top_n=3,
),
)
You can now interact with Pinecone's Inference API without the need to install any extra plugins.
Note: If you were previously using the preview versions of this functionality via the pinecone-plugin-inference package, you will need to uninstall that package.
from pinecone import Pinecone
pc = Pinecone(api_key="<<PINECONE_API_KEY>>")
inputs = ["Who created the first computer?"]
outputs = pc.inference.embed(
model="multilingual-e5-large",
inputs=inputs, parameters={"input_type": "passage", "truncate": "END"}
)
print(outputs)
# EmbeddingsList(
# model='multilingual-e5-large',
# data=[
# {'values': [0.1, ...., 0.2]},
# ],
# usage={'total_tokens': 6}
# )
asyncioThe v6 Python SDK introduces a new client variants, PineconeAsyncio and IndexAsyncio, which provide async methods for use with asyncio. This should unblock those who wish to use Pinecone with modern async web frameworks such as FastAPI, Quart, Sanic, etc. Those trying to onboard to Pinecone and upsert large amounts of data should significantly benefit from the efficiency of running many upserts in parallel.
To use these, you will need to install pinecone[asyncio] which pulls in an extra depdency on aiohttp. See notes on installation.
You can expect more documentation and information on how to use these asyncio clients to follow soon.
import asyncio
from pinecone import (
PineconeAsyncio,
IndexEmbed,
CloudProvider,
AwsRegion,
EmbedModel
)
async def main():
async with PineconeAsyncio() as pc:
if not await pc.has_index(index_name):
desc = await pc.create_index_for_model(
name="book-search",
cloud=CloudProvider.AWS,
region=AwsRegion.US_EAST_1,
embed=IndexEmbed(
model=EmbedModel.Multilingual_E5_Large,
metric="cosine",
field_map={
"text": "description",
},
)
)
asyncio.run(main())
Interactions with a deployed index are done via IndexAsyncio class, which can be instantiated using helper methods on either Pinecone or PineconeAsyncio:
import asyncio
from pinecone import Pinecone
async def main():
pc = Pinecone(api_key='<<PINECONE_API_KEY>>')
async with pc.IndexAsyncio(host="book-search-dojoi3u.svc.aped-4627-b74a.pinecone.io") as idx:
await idx.upsert_records(
namespace="books-records",
records=[
{
"id": "1",
"title": "The Great Gatsby",
"author": "F. Scott Fitzgerald",
"description": "The story of the mysteriously wealthy Jay Gatsby and his love for the beautiful Daisy Buchanan.",
"year": 1925,
},
{
"id": "2",
"title": "To Kill a Mockingbird",
"author": "Harper Lee",
"description": "A young girl comes of age in the segregated American South and witnesses her father's courageous defense of an innocent black man.",
"year": 1960,
},
{
"id": "3",
"title": "1984",
"author": "George Orwell",
"description": "In a dystopian future, a totalitarian regime exercises absolute control through pervasive surveillance and propaganda.",
"year": 1949,
},
]
)
asyncio.run(main())
Tags are key-value pairs you can attach to indexes to better understand, organize, and identify your resources. Tags are flexible and can be tailored to your needs, but some common use cases for them might be to label an index with the relevant deployment environment, application, team, or owner.
Tags can be set during index creation by passing an optional dictionary with the tags keyword argument to the create_index and create_index_for_model methods. Here's an example demonstrating how tags can be passed to create_index.
from pinecone import (
Pinecone,
ServerlessSpec,
CloudProvider,
GcpRegion,
Metric
)
pc = Pinecone(api_key='<<PINECONE_API_KEY>>')
pc.create_index(
name='my-index',
dimension=1536,
metric=Metric.COSINE,
spec=ServerlessSpec(
cloud=CloudProvider.GCP,
region=GcpRegion.US_CENTRAL1
),
tags={
"environment": "testing",
"owner": "jsmith",
}
)
See this page for more documentation about how to add, modify, or remove tags.
Sparse indexes are currently in early access. This release will allow those with early access to create sparse indexes and view those configurations with the describe_index and list_indexes methods.
These are created using the same create_index method as other index types but with different configuration options. For sparse indexes, you must omit dimension while passingmetric="dotproduct" and vector_type="sparse".
from pinecone import (
Pinecone,
ServerlessSpec,
CloudProvider,
AwsRegion,
Metric,
VectorType
)
pc = Pinecone()
pc.create_index(
name='sparse-index',
metric=Metric.DOTPRODUCT,
spec=ServerlessSpec(
cloud=CloudProvider.AWS,
region=AwsRegion.US_WEST_2
),
vector_type=VectorType.SPARSE
)
# Check the description to get the host url
desc = pc.describe_index(name='sparse-index')
# Instantiate the index client
sparse_index = pc.Index(host=desc.host)
Upserting and querying a sparse index is very similar to before, except now the values field of a Vector (used when working with dense values) may be unset.
import random
from pinecone import Vector, SparseValues
def unique_random_integers(n, range_start, range_end):
if n > (range_end - range_start + 1):
raise ValueError("Range too small for the requested number of unique integers")
return random.sample(range(range_start, range_end + 1), n)
# Generate some random sparse vectors
sparse_index.upsert(
vectors=[
Vector(
id=str(i),
sparse_values=SparseValues(
indices=unique_random_integers(10, 0, 10000),
values=[random.random() for j in range(10)]
)
) for i in range(10000)
],
batch_size=500,
)
# Querying sparse
sparse_index.query(
top_k=10,
sparse_vector={"indices":[1,2,3,4,5], "values": [random.random()]*5}
)
Many enum objects have been added to help with the discoverability of some configuration options. Type hints in your editor will now suggest enums such as Metric, AwsRegion, GcpRegion, PodType, EmbedModel, RerankModel and more to help you quickly get going without having to go looking for documentation examples. This is a backwards compatible change and you should still be able to pass string values for fields exactly as before if you have preexisting code.
For example, code like this
from pinecone import Pinecone, ServerlessIndex
pc = Pinecone()
pc.create_index(
name='my-index',
dimension=1536,
metric='cosine',
spec=ServerlessSpec(cloud='aws', region='us-west-2'),
vector_type='dense'
)
Can now be written as
from pinecone import (
Pinecone,
ServerlessSpec,
CloudProvider,
AwsRegion,
Metric,
VectorType
)
pc = Pinecone()
pc.create_index(
name='my-index',
dimension=1536,
metric=Metric.COSINE,
spec=ServerlessSpec(
cloud=CloudProvider.AWS,
region=AwsRegion.US_WEST_2
),
vector_type=VectorType.DENSE
)
Both ways of working are equally valid. Some may prefer the more concise nature of passing simple string values, but others may prefer the support your editor gives you to tab complete when working with enums.
pinecone-plugin-records and pinecone-plugin-inference are no longer needed because the functionality has been incorporated into the pinecone package itself. If you attempt to load the v6 client with these plugins present in your development environment, you will see an exception directing you to uninstall those plugins.tqdm which is used to provide a nice progress bar when upserting lots of data into Pinecone. If tqdm is available in the environment the Pinecone SDK will detect and use it but we will no longer require tqdm to be installed in order to run the SDK. Popular notebook platforms such as Jupyter and Google Colab already include tqdm in the environment by default so for many users this will not require any changes, but if you are running small scripts in other environments and want to continue seeing the progress bars you will need to separately install the tqdm package.config, openapi_config, and index_api) to instead prefer dedicated keyword arguments for individual settings such as api_key, proxy_url, etc. These keyword arguments were primarily aimed at facilitating testing but were never documented for the end-user so we expect few people to be impacted by the change. Having multiple ways of passing in the same configuration values was adding significant amounts of complexity to argument validation, testing, and documentation that wasn't really being repaid by significant ease of use, so we've removed those options.Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
This release contains a small adjustment to the query_namespaces method added in the 5.4.0. The initial implementation had a bug that meant it could n
This release contains a small adjustment to the query_namespaces method added in the 5.4.0. The initial implementation had a bug that meant it could not properly merge small result sets across multiple namespaces. This release adds a required keyword argument, metric to the query_namespaces method, which should enable the SDK to merge results no matter how many results are returned.
from pinecone import Pinecone
pc = Pinecone(api_key='YOUR_API_KEY')
index = pc.Index(host='your-index-host')
query_results = index.query_namespaces(
vector=[0.1, 0.2, ...], # The query vector, dimension should match your index
namespaces=['ns1', 'ns2', 'ns3'],
metric="cosine", # This is the new required keyword argument
include_values=False,
include_metadata=True,
filter={},
top_k=100,
)
Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v5.4.1...v5.4.2
[Chore] Allow support for pinecone-plugin-inference >=2.0.0, <4.0.0 by @austin-denoble in https://github.com/pinecone-io/pinecone-python-client/pull/4
pinecone-plugin-inference >=2.0.0, <4.0.0 by @austin-denoble in https://github.com/pinecone-io/pinecone-python-client/pull/419In this release we have added a utility method to run a query across multiple namespaces, then merge the result sets into a single ranked result set w
In this release we have added a utility method to run a query across multiple namespaces, then merge the result sets into a single ranked result set with the top_k most relevant results. The query_namespaces method accepts most of the same arguments as query with the addition of a required namespaces param.
Since query_namespaces executes multiple queries in parallel, in order to get good performance it is important to set values for the pool_threads and connection_pool_maxsize properties on the index client. The pool_threads setting is the number of threads available to execute requests while connection_pool_maxsize is the number of cached http connections that will be held. Since these tasks are not computationally heavy and are mainly i/o bound, it should be okay to have a high ratio of threads to cpus.
The combined results include the sum of all read unit usage used to perform the underlying queries for each namespace.
from pinecone import Pinecone
pc = Pinecone(api_key="key")
index = pc.Index(
name="index-name",
pool_threads=50, # <-- make sure to set these
connection_pool_maxsize=50, # <-- make sure to set these
)
query_vec = [ 0.1, ...] # an embedding vector with same dimension as the index
combined_results = index.query_namespaces(
vector=query_vec,
namespaces=['ns1', 'ns2', 'ns3', 'ns4'],
top_k=10,
include_values=False,
include_metadata=True,
filter={"genre": { "$eq": "comedy" }},
show_progress=False,
)
for scored_vec in combined_results.matches:
print(scored_vec)
print(combined_results.usage)
A version of query_namespaces is also available over grpc. For grpc, there is no need to set the connection_pool_maxsize because grpc makes efficient use of open connections by default.
from pinecone.grpc import PineconeGRPC
pc = PineconeGRPC(api_key="key")
index = pc.Index(
name="index-name",
pool_threads=50, # <-- make sure to set this
)
query_vec = [ 0.1, ...] # an embedding vector with same dimension as the index
combined_results = index.query_namespaces(
vector=query_vec,
namespaces=['ns1', 'ns2', 'ns3', 'ns4'],
top_k=10,
include_values=False,
include_metadata=True,
filter={"genre": { "$eq": "comedy" }},
show_progress=False,
)
for scored_vec in combined_results.matches:
print(scored_vec)
print(combined_results.usage)
query_namespaces by @jhamon in https://github.com/pinecone-io/pinecone-python-client/pull/409connection_pool_maxsize on Index and add docstrings by @jhamon in https://github.com/pinecone-io/pinecone-python-client/pull/415Full Changelog: https://github.com/pinecone-io/pinecone-python-client/compare/v5.3.1...v5.4.0.dev5
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →