NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #553 most downloaded on PyPI
Python bindings for CUDA
Last release 12 days ago
23 Sep 2026
Ships fairly regularly
a new release about every 6 weeks
Most releases are documented
notes for 21 of 25 stable releases
Nothing withdrawn
no release was ever pulled
2 years old
27 releases · first in 2025
One column per month.
Nothing published for this version
Nothing published for this version
Release notes: https://nvidia.github.io/cuda-python/cuda-bindings/13.4.1/release/13.4.1-notes.html
Release v13.4.1
Release notes: https://nvidia.github.io/cuda-python/cuda-bindings/13.4.1/release/13.4.1-notes.html
Refactor: tidy cdef declarations and struct init in .pyx files
Declare cdef locals at their point of initialization rather than at the
top of the function, and replace field-by-field struct setup with
Cython's struct-initializer syntax.
Only complete initializers are converted. Cython does not zero-fill
omitted members, so a partial initializer would leave them holding stack
garbage; sites that depend on a preceding memset are left unchanged.
Where every member is now supplied, the redundant memset is dropped.
No behavior change.
Add pre-commit check to keep pixi cuda-version pins in sync with ci/versions.yml (#2306)
Got initial version to address issue 2183. Let pre-commit check covers pixi cuda version pins to ci/versions.yml
rename to be more accurate
make error message more readable and accurate
put cuda_bindings / cuda_core to error message to be best accurate
add docstring
rename cuda_feature from cu13 to cu{major} to support bumping major version, e.g. 13.x.x to 14.x.x
add extracted line from pixi files to shown when check OK
add concrete build version alon side with expected version
polish to fix cosmetic
address pre commit check
Pin pyyaml in check-pixi-cuda-version pre-commit hook
coverage: add cuda.core tests for system device, checkpoint, memory, stream, green context, and tensor map (#2404)
Signed-off-by: Rui Luo ruluo@nvidia.com
test(core): add cuda.core.all vs public docs consistency check (#2347)
test(core): add cuda.core.all vs public docs consistency check
Closes #2326.
Parses docs/source/api.rst (autosummary entries and data directives while
cuda.core is the active module) and compares the flat public names against
cuda.core.all in both directions. Dotted entries such as graph.Graph or
checkpoint.Process are submodule namespaces and are excluded. Symbols
documented in api_private.rst are accepted as documented so returned-helper
docs do not fail the check.
The tests skip when cuda.core.all is not defined, so this lands
independently of #2300 and activates once #2300 merges. Also adds the
all-names-resolve guard suggested in the #2300 review.
Signed-off-by: Aryan aryansputta@gmail.com
Addresses review feedback that the consistency check was too narrow:
test(core): parse API docs with docutils
Update content to pass current tests
Add docutils dependency to pyproject.toml
Simplify checks. No longer make sure that everything documented is public.
test(core): document intentional scope limits of api docs consistency check
Reorganize all
Fix doc reference
Address findings in PR
Fix tests and make all construction consistent
test(core): drop IPC types from _memory package contents expectation
_ipc.all is now empty, so from cuda.core._memory import * no longer
binds IPCAllocationHandle or IPCBufferDescriptor. Update the expected list
in test_package_contents to match.
Signed-off-by: Aryan aryansputta@gmail.com
Signed-off-by: Aryan aryansputta@gmail.com
Co-authored-by: Michael Droettboom mdroettboom@nvidia.com
Co-authored-by: Michael Droettboom mdboom@gmail.com
Check for release notes on the tagged commit (#2446)
Fix version parsing in cuda_core; Allow new enums to be added to cuda_bindings (#2451)
Fix version parsing in cuda_core
Fix enum checks
chore: cython warnings as errors in cuda.core (#2441)
chore: -Werror for cythonization in cuda.core
chore: -Werror for cythonization in cuda.bindings
address review feedback
fix other cython warnings missed locally
[doc-only] docs: document setuptools-scm clone requirements for source builds (#2424)
docs: document setuptools-scm clone requirements for source builds
Signed-off-by: Bharat Raghunathan bharatrgatech@gmail.com
Signed-off-by: Bharat Raghunathan bharatrgatech@gmail.com
Signed-off-by: Bharat Raghunathan bharatrgatech@gmail.com
Add cuda-core 1.1.1 release notes (#2452)
Add griffe to CI (#2300)
cuda.core: add graph definition node updates (#2395)
feat(cuda.core): add event record node updates
Use the generic node setter with failure-atomic attachment replacement, establishing the shared path for definition-level parameter mutation.
test(cuda.core): remove unused graph update import
feat(cuda.core): add event wait and host node updates
Extend definition-level mutation to event waits and both Python and ctypes host callbacks while preserving old executable state and attachment ownership.
Report unsupported driver or binding versions before preparing mutation attachments or calling the generic node setter.
Allow partial memset parameter replacement while preserving graph-owned destination lifetimes and previously instantiated graph behavior.
Support partial copy parameter replacement while preserving independent source and destination ownership across graph instantiations.
Support independent launch configuration and argument replacement while requiring explicit arguments when changing kernels.
Replace embedded child hierarchies while preserving attachment metadata, invalidating stale views, and keeping existing executables independent.
Describe supported mutation methods, CUDA 12.2 requirements, and executable graph behavior in the API and release notes.
Use type-erased shared ownership for prepared child updates so extension loading does not depend on a hidden C++ deleter symbol.
Reflect shared ownership for the opaque child update transaction in the generated stub.
Make memcpy and memset mutation calls explicit and unambiguous before the public API freezes.
Preserve memory-node contexts, reject unsupported node forms, and fail clearly when child graph metadata cannot be updated.
Resolve the CUDA 13.2 graph parameter getter dynamically so CUDA 12 binding builds remain compilable.
Clarify parameter handling and cover host/device memory transitions while exposing context-sensitive test teardown for follow-up.
Cover the public invalid-node state to ensure parameter updates fail cleanly without restoring graph membership.
Defect 4 of #2388.
Signed-off-by: Aryan aryansputta@gmail.com
Add identity preserving pattern to critical_sections (#2453)
Add identity preserving pattern to critical_sections
this is annoying, but at least the library and probably kernel
attributes should be idempotent.
But critical sections can and will be released (similar to the GIL
although not sure what is more likely).
The important thing to note here is that the final attribute setting
section is self-contained and holds the lock (even if another thread
may have already set the attribute or still be executing the code
above).
As per review by Keith
Minimal thread-unsafe initialization order fixes
Make Windows pathfinder dynamic library searches architecture-aware (#2393)
Make Windows pathfinder searches architecture-aware
Avoid Windows architecture detection on Linux
Rename unsupported architecture error
Skip CUDA 12 wheel paths on Windows ARM64
Restore ARM64 cudla CTK search path
Group Windows search paths by architecture
Add architecture-specific Windows path tables
Make Windows CTK libnames architecture-aware
Add Windows architecture availability helper
Correct cuSPARSELt Windows ARM64 wheel path
Correct Windows CTK NVVM and CUPTI paths
Validate Windows NVVM binary architecture
Remove Windows architecture availability helper
Remove site-package catalog generation tool
Move pathfinder changes to 1.6.1 release notes
Remove pathfinder catalog generation tools
Restore site-packages collection scripts
Use platform-specific supported library names
Regenerate pathfinder 1.6.1 release notes
Remove redundant Windows libname consistency test
Mark agent-authored pathfinder tests
Clarify Windows binary architecture validation
Use all available dynamic library names
Test Windows site-package libraries by architecture
Require exactly one Windows architecture flag
Co-authored-by: Michael Wang isVoid@users.noreply.github.com
Co-authored-by: Ralf W. Grosse-Kunstleve rgrossekunst@nvidia.com
fix(cuda.core): surface real CUDA error in Device.set_current() (#2461)
cuda.core: fix host memcpy source pointer cast (#2476)
CI: avoid SIGPIPE when selecting core release tag (#2475)
Make pre-commit work on Windows (#2327)
Make pre-commit work on Windows
Update .pre-commit-config.yaml
Address some of the comments in the PR
Simplify type-checking
Address comments in PR
Simplifications
Fix simplifications
Fix type check
Add comment about stubgen-pyx issues
Update CONTRIBUTING.md
Co-authored-by: Ralf W. Grosse-Kunstleve rwgkio@gmail.com
Co-authored-by: Ralf W. Grosse-Kunstleve rwgkio@gmail.com
Co-authored-by: Ralf W. Grosse-Kunstleve rgrossekunst@nvidia.com
As Keith noted, this is needed for using the py_safe_call_once
definitions, Cython 3.2.5 changelog:
https://cython.readthedocs.io/en/latest/src/changes.html
(I guess the bump in the pre-commit is likely not strictly needed, but
there also were no stub changes.)
ci: drop custom NumPy and use beta 4 for Python 3.15 (#2411)
ci: drop custom NumPy builds for Python 3.15
Signed-off-by: tirthpatel90 tirthpatel5393@gmail.com
Signed-off-by: tirthpatel90 tirthpatel5393@gmail.com
Signed-off-by: tirthpatel90 tirthpatel5393@gmail.com
Fix NumPy version in pyproject.toml and try re-adding windows python 3.15
Bump cibuildwheel to 4.1.1 (which uses containers with 3.15 b4)
Revert "Use Python 3.15b2 for now until cibuildwheel is updated (#2433)"
This reverts commit 3ef82d6.
Add allow-prereleases to windows CI to try and run 3.15
Exclude ml-dtypes from windows (builds in 1 minute on linux so kept it)
Drop windows 3.15t again as psutil doesn't have free-threaded wheels
Signed-off-by: tirthpatel90 tirthpatel5393@gmail.com
Co-authored-by: Sebastian Berg sebastianb@nvidia.com
Remove tests involving NVLINK_MAX_LINKS (#2483)
Make temperature threshold checks forward-compatible (#2488)
Compare raw NVML device architecture values so architectures newer than the generated DeviceArch enum do not raise ValueError before the threshold query. Add regression coverage for an unrecognized architecture value.
fix(cuda.core): fall back to driver when nvJitLink < 12.3 is installed (#2409)
fix(cuda.core): fall back to driver when nvJitLink < 12.3 is installed
Stop probing nvJitLink availability via module.version(), which calls
the unversioned nvJitLinkVersion symbol missing in nvJitLink 12.0-12.2.
Use symbol pointer inspection via _nvjitlink_has_version_symbol()
instead, restoring cuda-core 0.6.0 fallback behavior.
Fixes #2408
Signed-off-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Cursor cursoragent@cursor.com
Add regression tests for Linker.which_backend() and
_decide_nvjitlink_or_driver() when the nvJitLinkVersion symbol is
missing (nvJitLink 12.0-12.2).
Related to #2408
Signed-off-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Cursor cursoragent@cursor.com
Document the #2408 regression fix in the cuda.core 1.2.0 release notes.
Related to #2408
Signed-off-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Cursor cursoragent@cursor.com
Address review feedback: keep the >=12.3 version-symbol check inside
_optional_cuda_import's probe so a missing nvJitLink dylib still falls
back to cuLink. Continue avoiding module.version(), which raises
FunctionNotFoundError on nvJitLink 12.0-12.2 (#2408).
Add coverage for missing-dylib fallback and a guard that the probe does
not call module.version().
Signed-off-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Cursor cursoragent@cursor.com
Address review feedback: drop the probe side-effect and catch
DynamicLibNotFoundError around _nvjitlink_has_version_symbol so missing
dylibs still fall back to cuLink. Keep avoiding module.version() for
nvJitLink <12.3 (#2408).
Mark newly added tests with agent_authored authorship markers.
Signed-off-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Cursor cursoragent@cursor.com
Signed-off-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Cursor cursoragent@cursor.com
Signed-off-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Omar Atie atiaomar1978-hub@users.noreply.github.com
Co-authored-by: Cursor cursoragent@cursor.com
Co-authored-by: Michael Wang 13521008+isVoid@users.noreply.github.com
Fix #2377: Reorganize the test helpers (where appropriate) to cuda_python_test_helpers (#2384)
Experiment: Install test_helpers as a package
Try something different in CI
Reorganize all the tests
Update a few more imports
docs: update nvshmem4py link to /api/latest/ path (#2492)
NVSHMEM docs now live under /nvshmem/api/latest/; the old unversioned
deep link 404s and breaks lychee on rendered docs.
Add mitigations for intermittent test failures with CUDA OOM (#2484)
Add context sync to teardown in init_cuda fixture
Cap memory pool size in some tests
cuda.core: return CUmodule via as_py in ObjectCode.get_module (#2481)
cuda.core: return CUmodule via as_py in ObjectCode.get_module
Use the shared handle export path for legacy CUmodule interop instead of
constructing driver.CUmodule directly.
Signed-off-by: Jinfeng jinfengl@nvidia.com
Satisfy cython-lint after switching get_module() to as_py().
Signed-off-by: Jinfeng jinfengl@nvidia.com
Route as_py(CUmodule) through as_intptr for consistency with other handle exports.
Signed-off-by: Jinfeng jinfengl@nvidia.com
cuda.core: resolve default-stream context per call instead of caching it (#2490)
cuda.core: resolve default-stream context per call instead of caching it
LEGACY_DEFAULT_STREAM and PER_THREAD_DEFAULT_STREAM wrap default-stream
tokens, which denote whatever context is current. Both are module-level
singletons, and Stream_ensure_ctx / Stream_ensure_ctx_device stored the
first context and device they observed on the object and never cleared
them, so a process-wide object became permanently bound to one context.
Replace the two helpers with resolvers that return the context and device
through out-parameters and cache on the object only when the stream is not
a default-stream token. Stream.context, .device, .resources, .record(), and
repr now follow the current context, a query no longer pins a context
reference for the lifetime of the process, and the shared singletons are no
longer written to from multiple threads.
Object identity is preserved, so eq and hash keying off the handle
are unaffected.
Fixes #2485
Skip sticky context reuse on default-stream tokens, document ambient
context behavior on device/resources/record, and cover resources in the
multi-GPU regression test.
Co-authored-by: Andy Jost ajost@nvidia.com
build: report internal build dependency provenance (#2509)
Fix nvbug6550424: Don't check NVML init behavior on CTK >= 13.4 (#2512)
cuda.core: add executable graph node updates (#2473)
Add executable graph attachment ownership
Install a private CUDA user object per graph executable so later node updates can retain replacement resources safely.
Expose ephemeral graph-node views that update complete executable parameters while retaining every replacement resource CUDA may still use.
Exercise public mutators, rollback, source reclamation, independent ownership, whole updates, and in-flight cleanup end to end.
Instantiation and whole-graph update each went through a prepare/commit
pair. That exposed an opaque transaction type over the internal C++
interface and split the exec ownership contract between C++ and Cython,
unlike every other resource handle, which a single create_* function
owns end to end.
Replace the pairs with create_graph_exec_handle and graph_exec_update.
Each stages a fresh attachment accumulator on the source graph, makes
the CUDA call with the GIL released, and adopts or publishes the result,
so the staging transaction becomes a stack guard in the anonymous
namespace instead of a header type. Cython keeps only what belongs to
it: filling the instantiation params and decoding the failure reasons.
The two driver entry points move into the C++ loader table with the
calls.
Convert the attachment append transaction to the unique_ptr plus
rollback deleter pattern that node attachments already use, which
retires the committed flag in favor of the same release-and-delete
mechanism. Drop GraphExecBox::attachment_object, which nothing reads.
Three gaps remained around owners attached to an executable graph.
Sequential updates to the same node must keep the superseded owner
reachable, because CUDA cannot detach user objects from an executable;
verified by breaking the append into a replace, which fails the new
test on exactly that assertion.
Closing an executable while a launch is in flight must not retire the
accumulator, since the launch still writes through the buffer that an
individual node update attached.
A child-graph update attaches no owner of its own and relies on CUDA
cloning the replacement graph's user object references into the
executable. Assert that contract directly: the callback outlives the
definition that supplied it and is released with the executable.
CUDA accepts user objects on a CUgraph only, so an executable graph can
never receive an owner after it exists. Document the consequence: one
accumulator is retained on the source graph, propagated by instantiation
or whole-graph update, and then released from the source so the
executable becomes its only owner.
Record why an owner is never removed once appended, and correct the two
Scope entries that still described executable graphs as untracked.
State the retention limit in the release notes as well. The API
reference already documents it, but the note is what a reader sees when
adopting the feature, and retention that looks unbounded deserves the
warning there.
Describe retention and complete-replacement rules without framing the
notes around CUDA limitations, and shorten the executable attachment
design section to problem, solution, and append-vs-replace limits.
Describe how callers use the view rather than how the binding retains
handles or when CUDA validates the node association.
Clarify attachment ownership and deferred-cleanup docs, trim the
invariants list to the cross-cutting rules, initialize instantiate
params with Cython struct syntax, and simplify the executable update
test helper.
The shared replacement fixture carries extra fields that are not update
parameters, so spreading it as kwargs breaks memset and kernel cases.
Add check and agent guidance about uncapped mempools (#2514)
Add check and agent guidance about uncapped mempools
Address review: move mempool check to pre-commit, share POOL_SIZE
Review feedback on #2514:
Resolve the CUDA path before importing cuda.bindings so the existing pathfinder import repairs PEP 517 namespace shadowing first. Reuse the resolved path for the CUDA include directory.
Make Windows static library searches architecture-aware (#2491)
Make static library discovery architecture-aware
Use architecture-specific static library paths
Co-authored-by: Michael Wang isVoid@users.noreply.github.com
cuda.core: add LaunchConfig.programmatic_stream_serialization for programmatic dependent launch (#2456)
feat(cuda.core): expose PDL via LaunchConfig.programmatic_stream_serialization
Allow users to set CU_LAUNCH_ATTRIBUTE_PROGRAMMATIC_STREAM_SERIALIZATION
through LaunchConfig, matching the is_cooperative attribute pattern (#1334).
Add an end-to-end Hopper+ test that launches primary and secondary kernels
on the same stream with programmatic_stream_serialization, and asserts
overlap only when the PDL attribute is enabled (#1334).
Drop unused secondary sync/sleep from the overlap test, clarify the
primary clock window comment, and print a short success line for CI.
add pre-commit passed
revise to xfail
Remove last remnants of Python 3.9 support (#2394)
Match private extension modules against the in-package path only (#2504)
check_cython_abi's private-module filter tested so_path.parts on the
absolute path, so any ancestor directory starting with an underscore made
every module look private. That is the normal layout under manylinux
(/opt/_internal/cpython-*/) and in GitHub Actions containers (/__w/), where
generate then writes zero ABI files and exits 0 -- a green run with no
coverage at all.
check's new-module scan had no filter, while generate skipped private
modules. Since generate never wrote an .abi.json for them, check reported
every private module as "New module added" on every run and set
has_allowed_changes, so it could not print "No changes found" for a package
shipping private submodules (cuda.bindings has _bindings/, _internal/, _lib/).
Extract the predicate into iter_public_extension_modules() so both paths use
it, and match only on the path relative to the package root.
Part 4 of the series proposed in #2410.
Filesystem predicates and path joining in the pathfinder tests now go through
pathlib: os.path.isfile/isdir become Path.is_file()/is_dir(), os.path.basename
becomes Path.name, os.path.join becomes Path joining, and the site-packages
check uses Path.parts instead of splitting on os.path.sep.
site_pkg_rel.replace("/", os.sep) is dropped in test_find_static_lib.py: Path
already accepts forward slashes on Windows.
Two files are left out on purpose. test_find_nvidia_binaries.py moves with
part 3, whose signature changes it depends on. test_search_steps.py is being
edited by #2489 (part 1), so converting it here would only create a conflict.
Left on the stdlib modules: glob.glob in test_find_nvidia_headers.py, which
expands an absolute pattern from the header catalog (Path.glob needs a base
dir, and the wildcard is not pinned to the last component); os.pathsep in
test_ctk_root_discovery.py, which builds PYTHONPATH, not a path; and os.sep in
test_utils_env_vars.py, which builds a trailing separator on purpose.
Signed-off-by: LeSingh1 sshaurya914@gmail.com
Windows coverage has not collected a test since 2026-03-17. The job builds
its wheels with a plain pip wheel, which never reads [tool.cibuildwheel],
so the delvewheel repair every other Windows build performs never ran here.
Those wheels import a bare "MSVCP140.dll" and resolve it against whatever the
test machine has in System32, which on the coverage runner is 14.00.24215.1,
built in 2015.
_resource_handles.pyd is compiled by MSVC 14.44 and imports exactly _Mtx_lock
and _Mtx_unlock from that DLL -- never _Mtx_init_in_situ, because std::mutex
has had a constexpr constructor since VS 2022 17.10. The 2015 runtime still
expects that initialisation and dereferences a null handle on the first lock,
which _stream.pyx takes while cuda.core is still importing. It is the only
extension module in either package that locks a mutex, which is why
cuda.bindings and cuda.pathfinder have always passed on the same machine.
Repairing the wheels vendors msvcp140 14.44 into cuda_core.libs and rewrites
the import tables to match, so the process no longer depends on what the test
machine carries. Verified on the coverage runner: 18 failed, 2929 passed,
918 skipped in 346s, against three to seven seconds of dying beforehand, and
the first Windows coverage data since March.
The same commit pins cuda-bindings to the wheel built one step earlier.
PIP_PRE is set so pip will consider that wheel at all -- it carries a .devN
version -- but it also admits PyPI's pre-releases, and cuda-bindings 13.4.0b1,
published 2026-07-29, outranks the local build. Its cydriver.pxd comes from
CTK 13.4 headers where CUmemLocation has a localized field, while cuda.core
compiles against the 13.3.0 mini-CTK where it does not, so the build has been
failing on error C2039 ever since.
Signed-off-by: Rui Luo ruluo@nvidia.com
cuda.core: accept ProgramOptions(name=None) (#2517)
cuda.core: accept ProgramOptions(name=None)
ProgramOptions.name is annotated str | None, but post_init called
.encode() on it unconditionally, so passing None raised AttributeError
before any CUDA call was reached.
Normalize None to the documented default, matching how arch is handled
in the same method. The encoded value is identical to the existing
default path, so the bytes passed to nvrtcCreateProgram are unchanged.
Signed-off-by: Aryan aryansputta@gmail.com
Extend coverage past ProgramOptions construction to assert the
normalized name reaches ObjectCode.name, matching the shape of
test_program_compile_valid_target_type.
Signed-off-by: Aryan aryansputta@gmail.com
ObjectCode.name receives an already-normalized options.name, so the
compile-level assertion could not fail independently of the options
test. Inline the default literal instead of a module constant, which
kept a private symbol out of the generated stub.
Signed-off-by: Aryan aryansputta@gmail.com
Signed-off-by: Aryan aryansputta@gmail.com
Co-authored-by: Michael Droettboom mdboom@gmail.com
cuda.bindings: Fix API status handling (#2530)
cuda.bindings tests: enable BAR lookup test on GH200
cuda_bindings fixes
Co-authored-by: Ralf Juengling rjuengling@utskinnyjoe-dvt-65.ipp2u1.colossus.nvidia.com
cuda.core: validate pinned host memory pool support (#2487)
cuda.core: validate pinned pool support
Reject unsupported host memory pools during allocation instead of allowing a later copy to fail with CUDA_ERROR_INVALID_VALUE.
Signed-off-by: Uday Arora udaya@nvidia.com
Drop the unnecessary CUDA 12 fence around host_memory_pools_supported,
raise RuntimeError instead of a synthetic CUDAError, and keep the
regression test hardware-gated for devices without host memory pools.
Signed-off-by: Uday Arora udaya@nvidia.com
Co-authored-by: Andy Jost ajost@nvidia.com
VirtualMemoryResource.__init__ classifies "host", "host_numa" and
"host_numa_current" all as host-located (it clears self.device for
each), but is_host_accessible compared with == "host". A resource
configured with location_type="host_numa" or "host_numa_current"
therefore reported is_host_accessible is False and
is_device_accessible is False -- an impossible answer that propagates
to Buffer.is_host_accessible, which forwards to the memory resource.
Share a single _HOST_LOCATION_TYPES set between the constructor and
the property so the two classifications cannot drift again.
cuda.core: validate ctypes host callback signatures against CUhostFn (#2525)
cuda.core: validate ctypes host callback signatures against CUhostFn
Reject incompatible ctypes prototypes before CUDA sees them, document
the required ABI, and note the stronger checking in the 1.2.0 release notes.
Use getattr for private ctypes calling-convention constants so the
regenerated _host_callback.pyi type-checks cleanly.
The previous check inspected ctypes' private flags bits to identify the
calling convention. That is wrong on Windows: CPython defines
FUNCFLAG_STDCALL as 0, so a bitwise test can never match WINFUNCTYPE, and
every win-64 test job rejected a valid callback. The 0x2 fallback used when
_ctypes.FUNCFLAG_STDCALL is absent is FUNCFLAG_HRESULT, not stdcall.
Drop the calling-convention check rather than repair the bit arithmetic.
ctypes only honors stdcall when building a callback on 32-bit x86 Windows,
which cuda.core does not support, and FUNCFLAG_PYTHONAPI is never consulted
on the callback path, so CFUNCTYPE, WINFUNCTYPE, and PYFUNCTYPE all yield the
same FFI_DEFAULT_ABI thunk. That leaves the declared result and argument
types, which are reachable through the public restype/argtypes attributes.
Reading those public attributes also lets a function pointer taken from a
shared library be accepted once its restype and argtypes are declared, which
the class-level lookup could never see.
chore: avoid some warnings when running cuda.core tests (#2515)
ci: constrain internal builds to exact local wheels (#2510)
ci: constrain internal builds to exact local wheels
ci: keep CI tool tests in nightly workflow
ci: generate local wheel constraints in workflows
fix(pathfinder): place Windows arm64 cudart test fixtures under bin/arm64 (#2528)
fix(pixi): restore conda test deps and relock after #2384 (#2532)
#2384 inserted a pypi-dependencies header mid-table, moving conda test
deps to PyPI without updating lockfiles. Fresh CI installs then dropped
the local cuda-bindings/cuda-core source packages, causing ModuleNotFoundError.
ci: limit pytest duration reports to the slowest 20 tests (#2523)
ci: limit pytest duration reports to the slowest 20 tests
--durations=0 prints every test phase and floods CI logs. Report only
the slowest 20 instead.
Move the duration limit into pytest addopts so CI and local runs
share one default, instead of repeating --durations on every command.
Fix SM resource alignment discovery test (#2389)
Fix SM resource alignment discovery test
Refine SM discovery alignment coverage
Document CUDA 13.4 SM discovery behavior
feat(security): onboard security-suite (secret + CodeQL) scanning. (#2589)
Call the centrally maintained NVIDIA/security-workflows security suite rather
than wiring each scan separately: one pinned reference runs the Pulse secret
scan and CodeQL SAST, both explicitly enabled.
Replace .github/workflows/codeql.yml with the suite's SAST scan. Both publish
code scanning results under the category /language:python, so keeping the local
workflow would put two analyses on every commit that overwrite each other's
alerts. The suite performs the same analysis: python, build-mode none,
security-extended queries, on ubuntu-latest.
fix(cuda.core): avoid truncating graph queries (#2587)
fix(cuda.core): avoid truncating graph queries
perf(cuda.core): retain adjacency stack buffer
test(cuda.core): cover large predecessor graph queries
Verify exact edge identities so graph query regressions cannot pass through count-only checks.
Co-authored-by: Andy Jost ajost@nvidia.com
ci: add selective wheel build plumbing (#2464)
Fix Windows binary utility discovery on Arm64 (#2586)
Fix Windows binary utility discovery on Arm64
Clarify binary utility search order
Expand standalone installation documentation
Align standalone search step comments
Preserve literal Nsight launcher lookup
Cover Windows binary discovery fallbacks
Document Windows architecture selection
Harden Windows Arm64 utility discovery
Fix Windows pre-commit checks
Fix CUDA path precedence documentation
Document Windows binary utility discovery
Co-authored-by: Michael Wang isVoid@users.noreply.github.com
Co-authored-by: Ralf W. Grosse-Kunstleve rgrossekunst@nvidia.com
Use pathlib in cuda.pathfinder._static_libs (part 2 of #2410) (#2493)
Migrate _static_libs finders from os.path to pathlib
Part 2 of the series proposed in #2410, following the same conversion
style as part 1 (#2489).
Path construction, joining, and filesystem predicates in
find_static_lib.py and find_bitcode_lib.py now go through pathlib.Path
instead of os.path string manipulation. Both modules keep importing os
solely for os.environ.get("CONDA_PREFIX").
Compatibility is preserved: every entry point still accepts str, and
every function that documents or returns str still returns str. Path is
used strictly as the internal representation and converted back with
str() at each return, so LocatedStaticLib.abs_path, LocatedBitcodeLib
.abs_path, find_static_lib() and find_bitcode_lib() are unchanged in
both type and value. No signature changes.
Signed-off-by: LeSingh1 sshaurya914@gmail.com
Follow-up to the review feedback on #2489: the str-compatibility constraint
applies only to the public API.
The try_* methods and _no_such_file_in_dir now work in Path throughout. str()
is applied once, where abs_path is stored on the public LocatedStaticLib and
LocatedBitcodeLib. The relative-path constants go from os.path.join(...) to
forward-slash literals, matching how site_packages_dirs is already written in
the same dicts; Path normalizes the separator on Windows.
One behavior change: a CUDA_PATH or CONDA_PREFIX containing redundant
separators ("//", "/.") now produces a normalized abs_path, because Path
collapses them. Differential fuzzing against the pre-revision code (16k lookups
over randomized trees, comparing located paths and full error text) shows no
other difference, and none at all when those variables are free of redundant
separators.
Signed-off-by: LeSingh1 sshaurya914@gmail.com
Signed-off-by: LeSingh1 sshaurya914@gmail.com
Co-authored-by: Michael Droettboom mdboom@gmail.com
chore: fix Apache-2.0 license notice and attribution gaps (#2605)
chore: fix Apache-2.0 license notice and attribution gaps
An open-source license review flagged several Apache-2.0 compliance gaps.
This addresses three of them, plus the guard that let one class of them
through. Licensing metadata only; no logic changes.
Copyright notices (15 files)
Two different defects that happened to share a symptom:
Header guard (toolshed/check_spdx.py)
COPYRIGHT_REGEX made "& AFFILIATES. All rights reserved." optional, so a
bare "NVIDIA CORPORATION" satisfied pre-commit. The suffix is now
required. (The 14 example files were passing for a different reason:
.spdx-ignore excludes cuda_bindings/examples/ entirely. That exclusion is
left alone here, but the files now conform, so it can be dropped in a
follow-up if desired.)
Tightening the regex surfaced two pre-existing files whose notice was
split or truncated -- cuda_core/cuda/core/_include/layout.hpp and
toolshed/build_static_bitcode_input.py. Both are corrected so the
mandated sentence appears verbatim on one line.
Third-party attribution (cuda_core/NOTICE)
cuda/core/_include/aoti_shim.h is a vendored subset of PyTorch's AOT
Inductor stable C ABI, BSD-3-Clause, carrying the upstream Facebook,
Idiap, Deepmind, NEC and NYU copyright lines, but NOTICE listed only
DLPack. A PyTorch entry is added with the full copyright block. The
accompanying aoti_shim.def carries no copyright line of its own and is
covered explicitly by that entry rather than given an NVIDIA header,
since it declares the same upstream symbol names. The DLPack entry now
also records where it is vendored.
LICENSE files (all five)
Every LICENSE ended at "END OF TERMS AND CONDITIONS", omitting the
required "APPENDIX: How to apply the Apache License to your work" and
its boilerplate. Appended to all five. The text is verified identical
to the canonical Apache 2.0 appendix.
Verified: 0 files with a non-conforming copyright string; check_spdx.py
passes over all 868 in-scope tracked files with the tightened regex.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Signed-off-by: Rob Parolin rparolin@nvidia.com
OSRB (NVBUG 4707569, comment #22) flagged the four sub-component LICENSE
files as redundant with the root LICENSE and asked for either their removal
or a root README Licensing section naming each subproject, its license and
its license path.
Each subproject builds an independent wheel and resolves its license file
relative to its own root, so the copies are kept and documented instead of
removed. Verified that the copies reach the built wheels: building
cuda_pathfinder produces dist-info/licenses/LICENSE even though its
pyproject.toml declares no explicit license-files (setuptools' default
LICEN[CS]E* glob covers it), as is also the case for cuda_core.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Signed-off-by: Rob Parolin rparolin@nvidia.com
Co-authored-by: Claude Opus 5 (1M context) noreply@anthropic.com
Use pathlib in toolshed and ci helper scripts (part 7 of
Note truncated.
Tag the 13.4.0b1 release
Tag the 13.4.0b1 release
https://nvidia.github.io/cuda-python/13.3.1/release/13.3.1-notes.html
conda install conda-forge::cuda-python=13.3.1conda install conda-forge::cuda-bindings=13.3.1conda install conda-forge::cuda-pathfinder=1.5.5https://nvidia.github.io/cuda-python/13.3.0/release/13.3.0-notes.html
conda install conda-forge::cuda-python=13.3.0conda install conda-forge::cuda-bindings=13.3.0conda install conda-forge::cuda-pathfinder=1.5.5https://nvidia.github.io/cuda-python/13.2.0/release/13.2.0-notes.html
https://nvidia.github.io/cuda-python/cuda-bindings/latest/release/13.1.1-notes.html
https://nvidia.github.io/cuda-python/cuda-bindings/latest/release/13.1.0-notes.html
https://nvidia.github.io/cuda-python/cuda-bindings/latest/release/13.0.3-notes.html
https://nvidia.github.io/cuda-python/cuda-bindings/latest/release/13.0.2-notes.html
https://nvidia.github.io/cuda-python/13.0.1/release/13.0.1-notes.html
https://nvidia.github.io/cuda-python/13.0.0/release/13.0.0-notes.html
Nothing published for this version
Nothing published for this version
https://nvidia.github.io/cuda-python/12.9.7/release/12.9.7-notes.html
conda install conda-forge::cuda-python=12.9.7conda install conda-forge::cuda-bindings=12.9.7conda install conda-forge::cuda-pathfinder=1.5.5https://nvidia.github.io/cuda-python/latest/release/12.9.6-notes.html
https://nvidia.github.io/cuda-python/latest/release/12.9.5-notes.html
https://nvidia.github.io/cuda-python/cuda-bindings/latest/release/12.9.4-notes.html
https://nvidia.github.io/cuda-python/cuda-bindings/latest/release/12.9.3-notes.html
https://nvidia.github.io/cuda-python/latest/release/12.9.2-notes.html
https://nvidia.github.io/cuda-python/latest/release/12.9.1-notes.html
https://nvidia.github.io/cuda-python/12.9.0/release/12.9.0-notes.html
https://nvidia.github.io/cuda-python/12.8.0/release/12.8.0-notes.html
https://nvidia.github.io/cuda-python/12.9.0/release/11.8.7-notes.html
https://nvidia.github.io/cuda-python/12.8.0/release/11.8.6-notes.html
Your coding agent can read these notes before it upgrades. Set up the MCP server →