NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2634 most downloaded on PyPI
Search, hash, sort, and process strings faster via SWAR and SIMD
Last release 2 days ago
02 Oct 2026
Release timing varies
gaps range from 8 days to 4 months
Most releases are documented
notes for 50 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
3 years old
110 releases · first in 2023
One column per quarter.
Add: Pyodide wheels and vectorized RISC-V builds
Release: v5.2.0 [skip ci]
Zvkg before dispatching the RISC-V GCM kernels (2f78f25)Fix: Silent offset wrap-around in arrow_strings_tape
Release: v5.1.2 [skip ci]
arrow_strings_tape (#331) (0d4d5c17)arrow_strings_tape::try_assign (#332) (013fe21d)sz_string_reserve when shrinking (#330) (9e096661)Release: v5.1.2 [skip ci]
arrow_strings_tape (#331) (0d4d5c1)arrow_strings_tape::try_assign (#332) (013fe21)sz_string_reserve when shrinking (#330) (9e09666)Fix: Snapshot caller sequences before walking them
Release: v5.1.1 [skip ci]
Release: v5.1.1 [skip ci]
AES-256 and SHA-256 are some of the world's most common workloads implemented in software and hardware to accelerate encryption and hashing. Both have
AES-256 and SHA-256 are some of the world's most common workloads implemented in software and hardware to accelerate encryption and hashing. Both have been optimized for decades and on x86 basically converged to more-or-less the same pure Asm snippets reused in most TLS & cryptography libraries. There are however multiple optimization axis still left to explore:
Those points come with their challenges and limitations. More on that and performance guidance numbers below.
The first axis is the only one SHA-256 has. One digest is a serial chain of 32 dependent SHA256RNDS2 at four cycles each, so no vector width shortens it, and every single-message implementation lands within a few percent of every other — they all ride the same instructions. Sixteen independent messages have no dependency between them at all, and there plain Skylake AVX-512 arithmetic pulls clear of the dedicated silicon:
| Hashing, one core | Throughput | Hardware |
|---|---|---|
sha2::Sha256 |
1.54 GB/s | 1 message, SHA-NI |
ring::SHA256 |
1.53 GB/s | 1 message, SHA-NI |
stringzilla::Sha256, single message |
1.41 GB/s | 1 message, SHA-NI |
stringzilla::Sha256s, dataset order |
0.90 GB/s | 16 lanes, AVX-512 |
stringzilla::Sha256s, length-sorted |
2.83 GB/s | 16 lanes, AVX-512 |
lanes = sz.Sha256s(16) # one hasher per lane
digests = bytearray(len(lanes) * sz.Sha256.digest_length)
lanes.update(batch).digest(out=digests) # no allocation, GIL released
tags = sz.hmac_sha256(b"secret", messages) # one tag per message, same kernels
The distance between the last two rows is the one limitation worth knowing. Lanes advance one block per turn, so a group costs as much as its longest member. They retire independently rather than dropping the whole group to scalar when one runs short, but XLSum lines run from 21 bytes to 445 KB, and in dataset order that skew leaves most slots idle. Feeding inputs in length order is the caller's lever, and argsort produces that ordering.
AES needs no help from the first axis, because counter mode is already parallel inside a single message. Four consecutive counter blocks of one stream encrypt at once, and GHASH follows once the key powers $H^1$ through $H^8$ are precomputed, so the problem reduces to how many blocks an implementation puts in flight per instruction. Profiling makes the ladder obvious, because everyone's hot symbol names its own width: Ring dispatches to aes_gcm_enc_update_vaes_avx2, the OpenSSL that Ubuntu 24.04 ships predates the vaes-avx512 assembly and runs AES-NI one block wide, and we issue _mm512_aesenc_epi128 with _mm512_clmulepi64_epi128.
| Encrypting, one core | Throughput | Width |
|---|---|---|
openssl::aes256gcm |
2.92 GB/s | 1-way, AES-NI |
ring::aes256gcm |
4.33 GB/s | 2-way, VAES on YMM |
stringzilla::aes256gcm |
6.41 GB/s | 4-way, VAES on ZMM |
stringzilla::aes256ctr |
9.23 GB/s | 4-way, VAES on ZMM |
Decryption tracks it. Counter mode pulls further ahead than the ladder alone predicts, because OpenSSL hand-wrote AVX-512 for GCM and the chaining modes but never for CTR.
key = sz.Aes256GcmKey(bytes(32))
ciphertext, tag = key.encrypt(b"hello", bytes(12))
assert key.decrypt(ciphertext, bytes(12), tag) == b"hello"
seekable = sz.Aes256CtrKey(bytes(32)) # unauthenticated, and its own inverse
assert seekable.xor(seekable.xor(b"hello", bytes(12)), bytes(12)) == b"hello"
Counter mode takes an absolute byte offset, so a reader after row group four hundred of a million-row file starts there rather than pushing the preceding gigabyte through the AES units. Galois/counter mode adds a tag and gives up seeking to get it, and streaming splits by type rather than by a flag — Aes256GcmEncryptor seals and Aes256GcmDecryptor opens, so a state pointed the wrong way is a compile error rather than a runtime one. Beyond x86 the backends cover Arm crypto extensions and SVE2-AES, RISC-V Zvkned with Zvkg, Power vcipher with vpmsumd, and a WebAssembly path with no cipher instructions at all, which emulates the round function through a constant-time tower-field substitution — the 256-byte lookup table a serial implementation reaches for is itself the side channel.
Both tables come from the StringWars harness over the same XLSum corpus, on one Intel Sapphire Rapids core, against OpenSSL 3.0.13 and Ring 0.17.14. Per-backend and per-size breakdowns live in
include/stringzilla/cipher/README.mdandinclude/stringzilla/hash/README.md.
evex512 version window at both ends (c437048e)const context (2b4b45c2)Improve: Report what the build is doing
Release: v5.0.7 [skip ci]
Make: Give nvcc on Windows an MSVC environment
Release: v5.0.6 [skip ci]
Release: v5.0.6 [skip ci]
Docs: Condense README tables into ASCII blocks
Release: v5.0.5 [skip ci]
__sz::split and the other 13 range helpers took the haystack by value__ — sz::split(view, "-") never compiled, and an owning lvalue was deep-copied so
sz::split and the other 13 range helpers took the haystack by value — sz::split(view, "-") never compiled, and an owning lvalue was deep-copied so every returned offset pointed into a private copy and was meaningless against your own iterators. Under SSO the offsets came back small and plausible rather than obviously wrong.sz_caps_none_k, so driver-less machines and containers reported no backends at all and silently lost every SIMD tier.sz_isascii helper is retired in favour of an ISA-specific byteset scan. The old SWAR loop could never reach the SIMD kernels; callers now name the backend explicitly and one that names none fails to compile instead of silently settling for the serial tier. 4.9 ns on NEON vs 7.0 ns for the retired SWAR over 100-byte inputs.cp314t) wheels, CUDA-on-Arm wheels building under Clang as host compiler, and the Swift package building on Windows.sz_isascii for a byteset scan (050d5f67)sz_rfind_byte_serial (0e21398c)cp314t wheels (3f430acb)sz::split and the other 13 range helpers took the haystack by value — sz::split(view, "-") never compiled, and an owning lvalue was deep-copied so every returned offset pointed into a private copy and was meaningless against your own iterators. Under SSO the offsets came back small and plausible rather than obviously wrong.sz_caps_none_k, so driver-less machines and containers reported no backends at all and silently lost every SIMD tier.sz_isascii helper is retired in favour of an ISA-specific byteset scan. The old SWAR loop could never reach the SIMD kernels; callers now name the backend explicitly and one that names none fails to compile instead of silently settling for the serial tier. 4.9 ns on NEON vs 7.0 ns for the retired SWAR over 100-byte inputs.cp314t) wheels, CUDA-on-Arm wheels building under Clang as host compiler, and the Swift package building on Windows.sz_isascii for a byteset scan (050d5f6)sz_rfind_byte_serial (0e21398)cp314t wheels (3f430ac)Improve: Close binding coverage gaps in Python & JavaScript
Release: v5.0.3 [skip ci]
Make: Ship npm as per-platform prebuilt packages
Release: v5.0.2 [skip ci]
Make: Publish the npm package via OIDC trusted publishing
Release: v5.0.1 [skip ci]
cccl (aa88432d)forkunion submodule (e83a55fb)_MSVC_LANG (657f1bfd)v4 vectorized sorting, hashing, and edit distances. v5 goes after the part everyone agreed was too irregular for SIMD — Unicode — plus the hardware an
v4 vectorized sorting, hashing, and edit distances. v5 goes after the part everyone agreed was too irregular for SIMD — Unicode — plus the hardware and packaging that make it portable.
<table> <thead> <tr> <th align="left">Domain</th> <th align="left">StringZilla v4</th> <th align="left">StringZilla v5</th> </tr> </thead> <tbody> <tr> <td><b>Unicode</b></td> <td>SIMD-vectorized case folding, uncased search</td> <td> <b>+</b> NFC/NFD/NFKC/NFKD normalization, <b>+</b> grapheme, <b>+</b> sentence, <b>+</b> word, and <b>+</b> line-break iterators</td> </tr> <tr> <td><b>Bioinformatics</b></td> <td>~700 GCUPS Levenshtein & ~10 GCUPS NW/SW on an H100</td> <td>up to <b>~6 TCUPS</b> Levenshtein and <b>~700 GCUPS</b> NW/SW on the same H100</td> </tr> <tr> <td><b>Portability</b></td> <td>x86 & Arm</td> <td> <b>+</b> WebAssembly SIMD, <b>+</b> RISC-V RVV, <b>+</b> LoongArch LASX, <b>+</b> IBM POWER VSX </td> </tr> <tr> <td><b>Hashing</b></td> <td>on par with xxHash & aHash</td> <td><b>+</b> up to <b>3x faster</b> multi-seed digests of one input — for Cuckoo tables, Count-Min & Bloom sketches</td> </tr> <tr> <td><b>Fingerprints</b></td> <td>streaming ~390 MB/s</td> <td>over <b>700 MB/s</b> on an H100 at the same embedding dimension — for large-scale MinHash content dedup</td> </tr> </tbody> </table>
strstr across C, C++, and Rust (7cd19532)sz_find_delimiter_utf8 UTF-8 delimiter scan (b5f70c33)Str and Strs headers (cb5e7966)sz_utf8_norm backends for every remaining ISA + differential fuzz (20a36674)sz_utf8_norm backends (Skylake baseline + Ice Lake override) (6767bea7)sz_utf8_norm with self-documenting tables (70e5d626)evex512 & popcnt in AVX-512 kernel attributes (124d7e93)sz_sequence_intersect to tolerate duplicates (74b20a85)file output for RISC-V in the cross-build verify step (097feb76)syscall() on RISC-V under a strict -std via _DEFAULT_SOURCE (4797f069)main* branches (c890d21a)dynamic-dispatch feature with compile-time SIMD dispatch (e856f1d5)cuda_status_t via memory to dodge NVCC struct-return bug (1d9651de)out= buffer for Strs.argsort (b6e1a4bc)Fix: Guard intersect & lookup on 32-bit / WebAssembly targets
Release: v4.6.3 [skip ci]
Improve: Support free-threaded Python 3.14t
szs.capabilities under QEMU (59a0556)package-lock.json in CI (20f3d26)StringZillaTests (5655376)-lcuda from stringzillas-cuda wheel link line (534d3f3)Make: Gitignore .claude/ per-project settings
Release: v4.6.1 [skip ci]
.claude/ per-project settings (37c2d8f)table_positions in sz_sequence_intersect_serial and _ice (087ad7f)sz_order_skylake tail for length-mismatched strings (a3926e2)sz_find_byte_serial (#308) (267c757)sz_find_skylake / sz_rfind_skylake tail for null-byte needles (#312) (7d72c96)New fast case-insensitive substring search path for Georgian 🇬🇪
inline static warnings with C++23 modules (#287) (374adbf)sz_find_byteset_haswell (#293) (7f2899a)sz_size_bit_ceil (4003057)Fix: Shared library function signatures
Below are the performance numbers comparing the search throughput of unique "word" tokens across various languages of the Leipzig Wikipedia Corpora fo
<img width="4650" height="2325" alt="search-utf8-example" src="https://github.com/user-attachments/assets/f4886da8-3590-4cc2-a666-368222dc0331" />
Below are the performance numbers comparing the search throughput of unique "word" tokens across various languages of the Leipzig Wikipedia Corpora for a case-insensitive substring search that respects all Unicode 17.0 case-folding rules. This is arguably the only library providing full Unicode spec compliance for search operations besides the PCRE2 library, which is often order(s) of magnitude slower than even our serial baseline due to the extreme complexity of combining a complete RegEx engine with Unicode compliance.
| Corpora Language | Script | Serial Baseline, GB/s | AVX-512 for Ice Lake+, GB/s | Speedup |
|---|---|---|---|---|
| Latin (Basic) | ||||
| 🇬🇧 English | Latin | 1.15 | 10.93 | 11.9× |
| 🇮🇹 Italian | Latin | 0.81 | 10.63 | 14.7× |
| 🇳🇱 Dutch | Latin | 0.85 | 10.91 | 13.3× |
| Latin (Extended) | ||||
| 🇩🇪 German | Latin+ß | 0.74 | 9.36 | 13.6× |
| 🇫🇷 French | Latin+Acc | 0.73 | 8.37 | 15.1× |
| 🇪🇸 Spanish | Latin+ñ | 0.99 | 8.86 | 10.8× |
| 🇵🇹 Portuguese | Latin+Acc | 0.77 | 9.58 | 14.3× |
| 🇵🇱 Polish | Latin+Ext | 0.62 | 7.51 | 14.2× |
| 🇨🇿 Czech | Latin+Háčky | 0.43 | 6.10 | 17.1× |
| 🇹🇷 Turkish | Latin+İ/ı | 0.81 | 6.78 | 11.7× |
| 🇻🇳 Vietnamese | Latin+Tones | 0.41 | 6.38 | 17.9× |
| Cyrillic | ||||
| 🇷🇺 Russian | Cyrillic | 0.54 | 3.41 | 10.6× |
| 🇺🇦 Ukrainian | Cyrillic | 0.56 | 4.03 | 10.6× |
| Greek | ||||
| 🇬🇷 Greek | Greek | 0.31 | 7.04 | 22.5× |
| Caucasian | ||||
| 🇦🇲 Armenian | Armenian | 0.34 | 4.18 | 17.5× |
| 🇬🇪 Georgian | Georgian | 0.65 | 10.56 | 24.2× |
| Semitic | ||||
| 🇮🇱 Hebrew | Hebrew | 0.65 | 9.52 | 13.7× |
| 🇸🇦 Arabic | Arabic | 1.17 | 9.85 | 9.8× |
| 🇮🇷 Persian | Arabic+Ext | 0.41 | 11.83 | 43.1× |
| Indic | ||||
| 🇮🇳 Hindi | Devanagari | 1.25 | 10.99 | 16.3× |
| 🇧🇩 Bengali | Bengali | 0.72 | 11.03 | 25.9× |
| 🇮🇳 Tamil | Tamil | 1.09 | 11.70 | 21.0× |
| CJK & East Asian | ||||
| 🇯🇵 Japanese | CJK+Kana | 0.52 | 11.56 | 26.7× |
| 🇰🇷 Korean | Hangul | 2.98 | 11.58 | 3.5× |
| 🇨🇳 Chinese | CJK | 0.43 | 20.07 | 103.0× |
sz_utf8_case_agnostic API (a0507ee).empty() for small strings (ea258c1)span::operator==0 for new NVCC benchamrks (9b6911a)curl on Alpine for Rust kit (85af5b5)bash on Alpine for Rust toolchain (999ec64)NULL missing - use SZ_NULL (7e3dd35)sized_match_t constructor (916b23e)pending_idx in fast ASCII iterator (8614658)Strs.tape accessors (718f9c1)hmac_sha256 (29c5732)start/end for folded find (62ad6f7)Release: v4.4.2 [skip ci] ### Patch - Fix: Windows MSVC compilation
Release: v4.4.2 [skip ci]
Added sz_at_least(n) macro for C99's static array parameter syntax, enabling compile-time bounds checking on fixed-size array arguments. In C mode, Cl
Added sz_at_least(n) macro for C99's static array parameter syntax, enabling compile-time bounds checking on fixed-size array arguments. In C mode, Clang will now warn when passing undersized arrays to annotated functions. The macro expands to nothing in C++ for compatibility.
// Compiler can now warn if the digest buffer is smaller than 32 bytes
void sz_sha256_state_digest(..., sz_u8_t digest[sz_at_least(32)]);
// Lookup tables must be at least 256 bytes
void sz_lookup(..., char const lut[sz_at_least(256)]);
See LWN.net article for background on this feature and its use in the Linux kernel.
static n arrays (#289) (039c4b4)To my knowledge, this is the first ever properly vectorized case-folding (aka .to_lower()) implementation compliant with Unicode (v17) and using SIMD
To my knowledge, this is the first ever properly vectorized case-folding (aka .to_lower()) implementation compliant with Unicode (v17) and using SIMD (AVX-512 for Intel Ice Lake and newer). The results are remarkable across most languages, but it wasn't trivial to achieve. Unlike dense linear algebra workloads, such as in SimSIMD, no shared logic holds across all languages and code points here. After all, Unicode began in 1989 and covers languages and writing systems that took thousands of years to develop and decades to be organized into a standardized set of rules.
This implementation focuses on locale-independent conversion. It covers every one of 1000+ character folding rules in CaseFolding.txt of the Unicode spec, including:
memcpy-like paths for unicameral scripts, like Chinese, Japanese, and Korean.To benchmark all of those, I've extended the StringWars benchmarks with a new bench_unicode.rs and bench_unicode.py scripts and the bench_unicode.md report produced for two dozen datasets pulled from the Leipzig Wikipedia corpora. On most languages the performance is great, except for Georgian and Vietnamese for now:
| Language | Standard 🦀 | StringZilla 🦀 | Standard 🐍 | StringZilla 🐍 | ||
|---|---|---|---|---|---|---|
| English 🇬🇧 | 482 MB/s | 7.53 GB/s | 16x | 257 MB/s | 3.14 GB/s | 12x |
| German 🇩🇪 | 432 MB/s | 2.59 GB/s | 6x | 260 MB/s | 1.81 GB/s | 7x |
| Russian 🇷🇺 | 217 MB/s | 2.20 GB/s | 10x | 470 MB/s | 1.56 GB/s | 3x |
| French 🇫🇷 | 346 MB/s | 1.84 GB/s | 5x | 274 MB/s | 1.37 GB/s | 5x |
| Greek 🇬🇷 | 220 MB/s | 1.00 GB/s | 5x | 431 MB/s | 779 MB/s | 2x |
| Armenian 🇦🇲 | 223 MB/s | 908 MB/s | 4x | 470 MB/s | 746 MB/s | 2x |
| Vietnamese 🇻🇳 | 265 MB/s | 352 MB/s | 1x | 340 MB/s | 291 MB/s | 1x |
| Arabic 🇸🇦 | 232 MB/s | 1004 MB/s | 4x | 467 MB/s | 1.80 GB/s | 4x |
| Bengali 🇧🇩 | 314 MB/s | 6.17 GB/s | 20x | 694 MB/s | 2.91 GB/s | 4x |
| Chinese 🇨🇳 | 325 MB/s | 1.21 GB/s | 4x | 697 MB/s | 886 MB/s | 1x |
| Czech 🇨🇿 | 322 MB/s | 827 MB/s | 3x | 292 MB/s | 688 MB/s | 2x |
| Dutch 🇳🇱 | 471 MB/s | 4.73 GB/s | 10x | 262 MB/s | 2.97 GB/s | 11x |
| Farsi 🇮🇷 | 235 MB/s | 858 MB/s | 4x | 475 MB/s | 1.42 GB/s | 3x |
| Georgian 🇬🇪 | 294 MB/s | 192 MB/s | 1x | 689 MB/s | 488 MB/s | 1x |
| Hebrew 🇮🇱 | 233 MB/s | 1.01 GB/s | 4x | 473 MB/s | 1.86 GB/s | 4x |
| Hindi 🇮🇳 | 293 MB/s | 6.32 GB/s | 22x | 682 MB/s | 3.14 GB/s | 5x |
| Italian 🇮🇹 | 439 MB/s | 2.29 GB/s | 5x | 268 MB/s | 1.93 GB/s | 7x |
| Japanese 🇯🇵 | 330 MB/s | 3.51 GB/s | 11x | 726 MB/s | 2.00 GB/s | 3x |
| Korean 🇰🇷 | 314 MB/s | 861 MB/s | 3x | 623 MB/s | 2.80 GB/s | 4x |
| Lithuanian 🇱🇹 | 352 MB/s | 864 MB/s | 2x | 274 MB/s | 728 MB/s | 3x |
| Polish 🇵🇱 | 364 MB/s | 939 MB/s | 3x | 277 MB/s | 786 MB/s | 3x |
| Portuguese 🇧🇷 | 395 MB/s | 2.38 GB/s | 6x | 270 MB/s | 1.79 GB/s | 7x |
| Spanish 🇪🇸 | 414 MB/s | 2.38 GB/s | 6x | 272 MB/s | 1.80 GB/s | 7x |
| Tamil 🇮🇳 | 306 MB/s | 6.05 GB/s | 20x | 712 MB/s | 3.03 GB/s | 4x |
| Turkish 🇹🇷 | 326 MB/s | 852 MB/s | 3x | 284 MB/s | 706 MB/s | 2x |
| Ukrainian 🇺🇦 | 217 MB/s | 2.09 GB/s | 10x | 476 MB/s | 1.58 GB/s | 3x |
For a complete comparison, go to StringWars 😉
__builtin missing on MSVC (fdc95f3)utf8_case* (44fbb92)Make: Deprecate current UTF-32 unpacking code
On AMD Zen5 Turin CPUs on different datasets, StringZilla provides the following throughput for splitting around whitespace and newline characters on 5 vastly different languages. Chinese and Korean texts, for example, are both made of mostly 3-byte letters, but Korean uses a lot of whitespace characters for syllable separation, while Chinese doesn't use any. French and English both use a lot of single-byte whitespace characters, but French uses many accented letters that are 2-byte long in UTF-8.
| Library | English | Chinese | Arabic | French | Korean |
|---|---|---|---|---|---|
| Split around 8 newline combinations: | |||||
stringzilla::utf8_newline_splits |
15.45 GiB/s | 16.65 GiB/s | 18.34 GiB/s | 14.52 GiB/s | 16.71 GiB/s |
stdlib::split(char::is_unicode_newline) |
1.90 GiB/s | 1.93 GiB/s | 1.82 GiB/s | 1.78 GiB/s | 1.81 GiB/s |
| Split around 25 whitespace characters: | |||||
stringzilla::utf8_whitespace_splits |
0.82 GiB/s | 2.40 GiB/s | 2.40 GiB/s | 0.92 GiB/s | 1.88 GiB/s |
stdlib::split(char::is_whitespace) |
0.77 GiB/s | 1.87 GiB/s | 1.04 GiB/s | 0.72 GiB/s | 0.98 GiB/s |
icu::WhiteSpace |
0.11 GiB/s | 0.16 GiB/s | 0.15 GiB/s | 0.12 GiB/s | 0.15 GiB/s |
On Apple M2 Pro:
| Library | English | Chinese | Arabic | French | Korean |
|---|---|---|---|---|---|
| Split around 8 newline combinations: | |||||
stringzilla::utf8_newline_splits |
5.69 GiB/s | 6.24 GiB/s | 6.58 GiB/s | 6.70 GiB/s | 6.29 GiB/s |
stdlib::split(char::is_unicode_newline) |
1.12 GiB/s | 1.11 GiB/s | 1.11 GiB/s | 1.11 GiB/s | 1.13 GiB/s |
| Split around 25 whitespace characters: | |||||
stringzilla::utf8_whitespace_splits |
0.57 GiB/s | 2.45 GiB/s | 1.18 GiB/s | 0.61 GiB/s | 0.92 GiB/s |
stdlib::split(char::is_whitespace) |
0.59 GiB/s | 1.16 GiB/s | 0.99 GiB/s | 0.63 GiB/s | 0.89 GiB/s |
icu::WhiteSpace |
0.10 GiB/s | 0.16 GiB/s | 0.14 GiB/s | 0.11 GiB/s | 0.14 GiB/s |
try_replace_all for Rust (35ed227)sz_utf8_unpack_upto64 for iterators (3ea1857)utf8.h for new valid and find_nth interfaces (e0465d5)SZ_ENFORCE_SVE_OVER_NEON=0 by default (da5687d)e280xx masks in SVE2 (5434ebf)i8 greater-than in AVX2 (dd4c4b0)uint8x16_t init (97cf851)svcompact_u8 in SVE2 (302af92)skip_empty arg for Python compatibility (0279383)hash.h (36fa527)svmatch-ing zero characters in SVE2 kernels (6f045aa)short implicit casts (00bacfc)utf8_count_neon w/out u64 unpacking in loop (b583fa8)no_std builds and doctests (bb699e9)sz_sha256_*_ice (2bceb8d)env fields for .vscode/tasks.json (dda7704)SZ_DEBUG=0 in CMake (febbdac)Fix: Missing bounds checks in Rust
Release: v4.2.3 [skip ci]
movemask bitsets (7c42b98)order array (32b6350)head_length is pre-decremented to zero (1c5c7e8)std::enable_if for non-STL builds (568d90c)Make: Linux cross-compile matching Release CI
Release: v4.2.2 [skip ci]
--sysroot cross-compile commands (579c82d)ARM64_CNTVCT on Windows (5e6777d)arch=armv8.2-a in pragmas (636147d)libc++ in LLVM builds on MacOS (1c8b29b)_M_ARM64=1 flags for MSVC (25311a6)winnt.h(169) C1189 error (8ef98a9)mrs w/out inline Asm on MSVC (d804c9f)--sysroot for "Cross Compile" builds (d3d901d)Make: Deprecate old cross-compilation scripts
Exposing SHA-256 to GoLang was tricky. Clang worked fine. GCC failed. It turned out that GCC was too shy about inlining my code, resulting in excessive stack space usage... Now, JavaScript, Swift, and GoLang bindings all support incremental SHA-256 procedures 🥳. Thanks to @MarekKnapek for reducing the stack memory usage of the serial SHA variant!
Moreover, thanks to @laurenspriem for highlighting the SIGILL when probing ID registers on older Arm CPUs. I've now guarded first mrs probes with signal handlers. Ugly solution, but it may work 😅 I've also improved the capability detection code on Arm-based Windows machines, using the OS-specific <processthreadsapi.h> functionality, so now not only pure NEON, but also NEON+SHA+AES kernels, should be dispatched just fine!
Thanks to @ashbob999, StringZilla is also getting more stable Windows builds and stringzilla_bare coverage in our CI 🦺
-pedantic for POSIX extensions (e99d557)-lpthread and pointer size (7722bb1)serialize_capability for Ice Lake on Clang (58f8cf9)capabilities in Arm macOS builds (511a09e)./scripts and StringWars (5af84dd)-D CMAKE_SYSROOT in cross-compiling CI (a26fc73)alloc warnings (4868d7f)mrs for avoid SIGILL on older Arm (d2f8e97)sz_checksum (97f9ecf)Capabilities to GoLang (5f2cc97)noescape/nocallback for stateful hashes (f8d321f)io.Writer & hash.Hash64 interface for Go (05f89ca)sz_dispatch_table_init for Go (5ff7ba1)C.sz_checksum (652735d)Improve: Deprecate SVE2 hashing
User-facing updates:
Implementation details:
sz_hash in NEON backendsz_copy in SVE backendsz_cap_goldmont_k capability! (f70e927)neon+sha new capability! (fcb68a4)bench_token (bb077da)hmac_sha256 APIs (bf1971e)Sha256 class for Python (6ae7b75)before-all for dnf on Fedora & apt on Debian (fc74452)bench_unary costs (d8d19ce)uint32x4_t on MSVC (3dce631)sz_copy_sve tail issue (2fa818d)<arm_neon_sve_bridge.h> (e5b4496)svlasta_u64(svpfalse_b()) UB (4833e83)Thanks to @Algunenano and the broader ClickHouse team for help, back-porting StringZilla kernels to older CPUs 🤗 With this release:
Thanks to @Algunenano and the broader ClickHouse team for help, back-porting StringZilla kernels to older CPUs 🤗 With this release:
VAES extensions only needed for Ice Lake and newer.VAES debuted in Ice Lake (8cef111)Improve: Faster unaligned loads in fingerprints
Release: v4.0.15 [skip ci]
Fix: ODR violation for raise in C++
Release: v4.0.14 [skip ci]
raise in C++ (#238) (6e826d7)Make: Pull CUDA within CIBW_BEFORE_ALL
Release: v4.0.13 [skip ci]
CIBW_BEFORE_ALL (6316330)This release fixes a critical bug where non-owning Strs slices incorrectly copied entire parent data during GPU memory allocation, instead of just the
This release fixes a critical bug where non-owning Strs slices incorrectly copied entire parent data during GPU memory allocation, instead of just the slice portion. The fix ensures proper Apache Arrow-compatible StringTape format handling with correct offset normalization for zero-copy operations. GPU memory management is now significantly more efficient, eliminating unnecessary re-allocations when data already resides in GPU memory through intelligent parent chain traversal.
A new stringzillas.to_device() function enables explicit GPU memory pre-allocation, useful for testing and performance optimization:
import stringzilla as sz
import stringzillas as szs
# Create strings and slices
strs = sz.Strs(["hello", "world", "test", "data"])
slice_view = strs[1:3] # Non-owning view of ["world", "test"]
# Pre-allocate on GPU (if available)
gpu_strs = szs.to_device(strs)
gpu_slice = szs.to_device(slice_view) # Correctly handles slice offsets
Cross-platform builds are now more stable with fixes for Windows ARM64 cross-compilation, ensuring mutually exclusive architecture flags prevent header conflicts. The CI/CD pipeline correctly generates stringzillas-cuda packages by properly propagating environment variables through cibuildwheel. Enhanced test coverage includes complex Unicode scenarios with RTL text, emoji sequences, and different normalization forms. Documentation has been extended with Rust examples showcasing zero-copy compute_into APIs using StringTape format.
to_device tests w/out GPUs found (d701592)SZ_TARGET into CIBW env (b222e82)realloc for on-GPU views (77e67cf)compute_into API with StringTape (f4ad81e)to_device(Strs) for unicode (f3c5357)to_device (c78cd21)Strs slicing as in StringTape (b4f8d12)Improve: Striping APIs for Python
Release: v4.0.11 [skip ci]
Fix: Propagate big-endian & SIMD flags
Release: v4.0.10 [skip ci]
Make: Stricter SZ_IS_64BIT checks
Release: v4.0.9 [skip ci]
SZ_IS_64BIT checks (0edabf5)__i386__ & _M_IX86 as 64-bit (4f3d8dc)Docs: Outdated algorithm details
Release: v4.0.8 [skip ci]
sdist (e1966de)SZ_TARGET.env for PyPI sdists (57822a4)compute_into (b507a8f)compute_into APIs for Rust (103019c)__description__s (1467001)before-test override (d19e2fa)sz PyTests in szs runs (8abcdab)RawParts (a9cfa22)Fix: Use unsigned offsets & smaller tapes
Release: v4.0.7 [skip ci]
Fix: Move macros for MSVC visibility
Release: v4.0.6 [skip ci]
include/ from Fork Union (948ae74)void specialization for enable_if (225d55a)__attribute__ with MSVC (f2170d8)Make: Override get_requires_for_build_wheel
Release: v4.0.2 [skip ci]
get_requires_for_build_wheel (c83f1d3)Make: Delay NumPy include path resolution
Release: v4.0.1 [skip ci]
Our old algorithm didn't perform any memory allocations and tried to fit too much into the provided buffers. The new breaking change in the API allows…
This PR entirely refactors the codebase, separating the single-header implementation into separate headers. Moreover, it brings faster kernels for:
And more community contributions:
Huge thanks to our partners at Nebius for their continued support and endless stream of GPU installations for the most demanding computational workloads in both AI and beyond!
Sadly, most modern software development tooling is subpar. VS Code is just as slow and unresponsive as the older Atom and the other web-based technologies, while LSP implementations for C++ are equally slow and completely mess up code highlighting for files over 5,000 Lines Of Code (LOCs). So, I've unbundled the single-header solution into multiple headers, similar to SimSIMD.
Also, similar to SimSIMD, CPU feature detection has been reworked to separate serial implementations, Haswell, Skylake, Ice Lake, NEON, and SVE.
Biology is set to be one of the driving forces of the 21st century, and biological DNA/RNA/protein data is already one of the fastest-growing data modalities, outpacing Moore's Law. Still, most of the BioInformatics software today is flawed and pretty slow. Last year, I helped several BioTech and Pharma companies scale up their data-processing capacity with specialized optimizations for various niche use cases. Still, I also wanted to include baseline kernels for the most crucial algorithms to StringZilla, covering:
These kernels are hardly state-of-the-art at this point, but should provide a good baseline, ensuring correctness and equivalent outputs across different CPU & GPU brands.
Our old algorithm didn't perform any memory allocations and tried to fit too much into the provided buffers. The new breaking change in the API allows passing a memory allocator, making the implementation more flexible. It now works fine on both 32-bit and 64-bit systems.
The new serial algorithm is often 5 times faster than the std::sort function in the C++ Standard Template Library for a vector of strings. It's also typically 10 times faster than the qsort_r function in the GNU C library. There are even quicker versions available for Ice Lake CPUs with AVX-512 and Arm CPUs with SVE.
Our old algorithm was a variation of the Karp-Rabin hash and was designed more for rolling hashing workloads. Sadly, such hashing schemes don't pass SMHasher and similar hash-testing suites, and a better solution was needed. For years, I have been contemplating designing a general-purpose hash function based on AES instructions, which have been implemented in hardware for several CPU generations now. As discussed with @jandrewrogers, and can be seen in his AquaHash project, those instructions provide an almost unique amount of mixing logic per CPU cycle of latency.
Many popular hash libraries, such as AHash in the Rust ecosystem, cleverly combine AES instructions with 8-bit shuffles and 64-bit additions. However, they rarely harness the full power of the CPU due to the constraints of Rust tooling and the complexity of using masked x86 AVX-512 and predicated Arm SVE2 instructions. StringZilla does that and ticks a few more boxes:
--extra tests.Implementing this logic, which provides both fast and high-quality hashes, often capable of computing four hashes simultaneously, made these kernels handy not only for hashing itself, but also for higher-level operations like database-style hash joins and set intersections, as well as advanced sequence alignment algorithms for bioinformatics.
Despite deprecating Rabin-Karp rolling hashes for general-purpose workloads, it's hard to argue with their usability in "fingerprinting" or "sketching" tasks, where a fixed-size feature vector is needed to compare contents of variable-length strings. The features should be as different as possible, covering various substring lengths, so polynomial rolling hashes fit nicely!
That said, implementing modulo-arithmetic over 64-bit integers is extremely expensive on Intel CPUs:
VPMULLQ (ZMM, ZMM, ZMM) for _mm512_mullo_epi64:
VPMULLD (ZMM, ZMM, ZMM) for _mm512_mullo_epi32:
VPMULLW (ZMM, ZMM, ZMM) for _mm512_mullo_epi16:
VPMADD52LUQ (ZMM, ZMM, ZMM) for _mm512_madd52lo_epu64 for 52-bit multiplication:
It's even pricier on Nvidia GPUs, but as you can see above, we may have a way out! x86 has a cheap instruction for 52-bit integer multiplication & addition. Moreover, 52 is precisely the number of bits we can safely use to exactly store an integer inside a double-precision floating-point number, opening doors to a really weird, but equally as implementation of rolling hash functions across CPUs and GPUs, implemented using 64-bit floats, and Barrett's reductions for modulo-arithmetics, to avoid division!
linear_score_on_each_cuda_warp_ and affine_score_on_each_cuda_warp_, which scale to 16 blocks of (228 - 1) KB shared memory on each, totaling 3.5 MB of shared memory that can be used for alignments. With u32 scores and 3 DP diagonals (in case of non-affine gaps) that should significantly accelerate the alignment of ~300K-long strings. But keep in mind that cluster.sync() is still very expensive - 1300 cycles - only 40% less than 2200 cycles for grid.sync().stringzilla_bare builds on MSVC, if needed.Strs ops (009b975)sz::edit_distance -> Levenshtein (d44beb4)lookup and fill_random (1ce830b)charset/generate -> byteset/fill_random (2ce2b49)similarity.h (095bc2d)look_up_transform to lookup API (e0055d5)checkum to bytesum, new hash, and PRNG (71f1f4b)sz_sort now takes allocators (ec81663)char_set constructor with literals (2c49eae)Str.count_byteset for Python (802d699)HashMap traits for Rust (e555cc3)try_resize_and_overwrite (e00a98b)sz.fill_random for Python (3565f9b)basic_rolling_hashers CUDA port (a25e3f2)LevenshteinDistancesUTF8 in Py (4cd6c7d)DeviceScope for Python (30b8fd3)Strs from lists, tuples, generators (05a5434)Strs layout conversion tests (4e5df62)_sz_py_api capsules (9d5935b)Strs.from_arrow conversion (0441b32)SZ_NOINLINE (4e6e1af)float fingerprints (3955eee)lock_guard to avoid STL (e6cdb93)to_span, to_view helpers (4fb9283)basic_rolling_hashers (63447d3)double fingerprinting (9ed6006)find_many minimal counter & benchmarks (3c615d9)bench_find_many.cpp (2d594e4)sz_similarity_gaps_t enums (5b410b8)safe_vector, safe_array (ac6218a)indexed_container_iterator (a5ca2ac)sz_caps (f76fa40)gpu_specs and cuda_status_t (c36d2b8)constant_iterator (6751e79)arrow_strings_tape::try_append (6d7f221)score_diagonally (427d5b5)lookup transform in Rust (3090472)find_byteset in Rust (667ea91)sz_find_sve (d7ede5d)find_byte with SVE (56a49c8)argsort, intersect (5ea0698)status_t for errors in C++ (63daa5f)sz_sequence_t helpers (dc7c109)sz_sequence_argsort_ice (69d4ecb)setuptools in CI (492ecc0)RuntimeError for engine calls (f9cbb00)itertools (f2704d7)Str_like_* naming convention (5fdd8ee).releaserc (8b5af6e)build.rs (d129f95).capabilities to JS (c557a62)basic_string (5fb0e74)__cpp_lib_string_resize_and_overwrite test guards (cb2d6a8)SZ_AVOID_STL (c3040b7)_MSC_VER to __GNUC__ conditions (83492ac)DeviceScope behavior (f217b86)s390x (8d2d9c8)E722 import exception warning (3aedfb2)np.random v1 vs v2 compatibility also in szs. (5f8ac03)np.random v1 vs v2 compatibility (aea096b)SZ_IS_QEMU_ (d433db0)sys.getrefcount tests on PyPy (80fa90e)-arch (6a450f6)pip install . (c462f42)+wfxt target (337257b)MACOSX_DEPLOYMENT_TARGET (95a95fe)fail-fast for Python pre-release wheels (a7c3f04)90a in Py & C (c317280)affine-gaps dep (fa52363)SWIFT_VERSION for container (a4015f1)capabilities_mode in PyTest (2bee1e4)bytes_per_cell <= 2 (48b4406)affine-gaps (48a538c)szs_capabilities (42f043c)/std:c++ for MSVC (930b1f0)bare builds on Windows & macOS (b42a340)sz before szs in CI (9033649)universal builds (ce6cab2)szs_ symbols (1fd5db8)osx (d06d4e4)SZ_DYNAMIC_DISPATCH check (acdab3f)wheel in CI (62f8c98)allocator_traits include (44b8d72)static_cast (b6a4cc6)static_cast to standard for MSVC (43c953f)uv in GitHub CI (89ed74d)group_by (3761107)qsort on MacOS (cd263ae)::capability_k (ddc640b)count (ef3ca96)levenshteinDistance in Swift (70c4add)capabilities from DeviceScope (a807eba)sz_device_scope_t (d9100a3)-O2 optimization (5fcde22)info1 (f4d4a76)sz_rune_parse (474dec4)value_type for CUDA fingerprinter (7724886)to_span compilation (e7fdd98)affine_levenshtein_utf8_ice_t (1e112c6)CudaBuildExtension for Python (4abf63c)prong_t in executor concepts (57d4ec8)LevenshteinDistancesUTF8Type (cb81c00)seeds in Strs.sample (5ac53c3)Strs from PyArrow (7aebef4)Strs offsets (65380ca)sz_capabilities in non-dynamic builds (6206eb4)rebind_alloc in C++20 (492d726)_mm256_cvtepi64_epi32 on Haswell (014002e)march flags through NVCC (496ae84)gpu_specs_fetch & GPU args order (3d7b491)lib.rs (5d82454)cuda_status_t (0a0955e)barrett_mod in C++ & Python (1fc1cab)uv for tests (f86dfaa)std::gcd (a1b3001)seed affects hashes (78d39f9)is_same_type usage over std::is_same (0ab5710)sz_bitcast strict aliasing (80b97de)std::swap in device code (c0aea26)std::remove_cvref to C++17 (ffee12b)safe_vector (78d8c96)constexpr use in C++11 (866e2f2)span extents (8355b6e)arrays_equality (e96d26f)find_many tests (c80ce60)+g with +m,r like GB (40bd3ed)count_many_parallel (0b22dd9)SZ_DYNAMIC attributes (2f0334a)nothrow-copyable views in tests (170b61b)bench_search target (1133cce)find_many.cuh compilation issues (72affc2)try_assign with different alloc (4f9d2db)fork_union (3f5de03)find_many algos for different length (0acbeeb)__reduce_max_sync in SW on Hopper (c45017d)unified_alloc propagation (13e0201)bench_find_many (16f5fa4)std::generate include (853a0fa)safe_vector (0c442d1)bench_search -> bench_find (6f65624)STRINGWARS_STRESS env-var (ade74f6)i16 SW on Hopper (817b15c)fork_union API in executors (1d6f58d)fork_union pools (81463d4)fork_union dependency (bd3d341)bytesum assignment (bfcd10e)views::group_by with callback (6a61e1b)bytes_per_cell_t enum (c6e907a)bench_unary (60b99ab)requires clause (22c691e)safe_vector (9c5a56c)std::execution for baseline tests (f228264)bytes_per_core_optimal estimate (83bc966)find_many_match_t properties (d0ebee8)sz_copy_skylake tail handling on large input (#222) (6da5e1e)i16 Ice Lake NW/SW alignment (8cc9794)similarity.hpp (10279f9)enums (35ba76c)constant_iterator in CUDA (bc59ee3)std::allocator::rebind deprecated (4c87404)std::iterator dependency (b3db596)sz_i16_t definition (238c86d)types.h (1c1582f)mean_token_length calculation (efcadd1)fsanitize (ea7647f)malloc(0) behavior (7524882)_mm256_set1_epi64x (aac2e8f)SZ_DEBUG macro (bca734a)sz::lookup examples in Rs (778d4f0)scripts/ deps for uv (1907d2b)uv instructions (9460fd4)sz_hash_state_t stores (7e65a1e)GiB over GB (811fc59)Byteset::from_bytes (34660f2)SZ_USE_SVE definitions (45b15b0)fill_random checksums in benchmarks (0bba772)find_sve mask update on long needles (a493ab8)bench_sequence CMake target (905749c)SQINCP in SVE for increments (efda23b)do_not_optimize token-level results (d1a3779)bench_sort (991a78b)std::search offsets (244e605)printf (298d214)std::string::data is mutable only since C++17 (ff23c3d)_sz_capabilities symbols (8bb90e5)sz_intersect signature (f712de3)stringzilla_bare on MacOS (a7b35ba)constexpr (feb415f)find_1byte signature compatibility (d19e8b8)sz_sequence_t::handle (75fabf1)cibuildwheel env variables (fbf256a)fill_random test condition (8b396c8)build.sh (2bbafa1)copy/move on Haswell with interleaving (69dfa10)sz_equal_haswell (7aad4bb)compare.h operations (d7bab8d)memory.h header (7698392)_sz_swap macro (dcf6c65)sz_sort to sz_qsort (6191cc6)sz_sort_serial passes tests (8bad799)uniform_int_distribution lower bound (bdee111)sz_sort_serial passes for same length inputs (0fda5a5)uniform_int_distribution upper bound (17f28a3)constexpr constructor must be empty (13bace2)std::accumulate for checksums (bce107a)checksum_haswell (b20d7cd)value_type (5bbd971)sz_checksum_haswell (84cb4c8)sz_checksum_haswell (509b58b)constexprs from C++20 to C++14 (0a3e363)usnigned (d18a159)levenshtein_baseline (d9557d3)BZHI (fa47deb)BZHI (bd7054e)sz_u512_vec_t members visibility (2007d49)stderr (084d653)basic_charset (864ee03)basic_charset operator (#203) (e20d207)stringzillite to stringzilla_bare (364e2ca)compare.h file (6512f1d)stringzilla.h file (41e5917)types.h file (b835051)sort.h file (1ba7982)small_string.h file (5b55e19)hash.h file (be4c63d)similarity.h file (8b401bd)memory.h file (295d49a)find.h file (2a1fcd1)sz_look_up_transform_avx512 declaration (585f7d5)#pragma region dashes (fe4449b)Fix: Skip second half of AVX2 copy
Release: v3.12.6 [skip ci]
Release: v3.12.5 [skip ci] ### Patch - Make: Upgrade to newer CMake
Release: v3.12.5 [skip ci]
Fix: Long input tail in sz_copy_avx512
Release: v3.12.4 [skip ci]
sz_copy_avx512 (#221) (18f04e7)Release: v3.12.3 [skip ci] ### Patch - Improve: C++ Lifetime bounds
Release: v3.12.3 [skip ci]
Release: v3.12.2 [skip ci] ### Patch - Fix: Horspool corner-case
Release: v3.12.2 [skip ci]
Fix: Uninitialized bad shift table
Release: v3.12.1 [skip ci]
Together with @MarkReedZ we've added basic GoLang bindings to StringZilla, which look surprisingly fast compared to native GoLang strings. We currentl
Together with @MarkReedZ we've added basic GoLang bindings to StringZilla, which look surprisingly fast compared to native GoLang strings. We currently use the new cGo annotations available in Go 1.24:
Cgo has gained new capabilities in Go 1.24, supporting new C function annotations to improve runtime performance. Among them,
#cgo noescape cFunctionNameis used to inform the compiler that the memory passed tocFunctionnamewill not escape;#cgo nocallback cFunctionNameindicates that this C function will not call back any Go functions. In addition, Cgo's inspection of multiple incompatible declarations of C functions has become more stringent. When there are incompatible declarations in different files, errors can be detected and reported more timely and accurately.
I was using an Intel Sapphire Rapids machine on AWS for preliminary testing and benchmarking. I've precompiled StringZilla with dynamic dispatch enabled, linked to the thin GoLang binding layer:
$ ~/StringZilla/golang$ CGO_CFLAGS="-I$(pwd)/../include" \
CGO_LDFLAGS="-L$(pwd)/../build_golang -lstringzilla_shared" \
LD_LIBRARY_PATH="$(pwd)/../build_golang:$LD_LIBRARY_PATH" \
go run ../scripts/bench.go --input ../leipzig1M.txt --split lines --seed 42
... and compared to native GoLang strings on some key operations:
Benchmarking on `../leipzig1M.txt` with seed 42.
Total input length: 129644797
Total lines: 1000000
Average line length: 128.64
Running benchmark using `testing.Benchmark`.
strings.Contains : 309 3818144 ns/op
sz.Contains : 664 1881251 ns/op
strings.Index : 325 3669081 ns/op
sz.Index : 624 1990093 ns/op
strings.LastIndex : 12 85201713 ns/op
sz.LastIndex : 494 2306318 ns/op
strings.IndexAny : 6321228 181.0 ns/op
sz.IndexAny : 10608960 112.6 ns/op
strings.Count : 156 8015292 ns/op
sz.Count (non-overlap) : 285 4206698 ns/op
sz.Count (overlap) : 284 4204370 ns/op
So if you are processing a lot of text in Go, try doing so with StringZilla and stay tuned for the upcoming 4.0 release #201 🥳
Release: v3.11.3 [skip ci] ### Patch - Improve: Pointer casting rules (a3f2f00) - Docs: LLVM build instruction
Release: v3.11.3 [skip ci]
Release: v3.11.2 [skip ci] ### Patch - Make: Using SIMD on FreeBSD
Release: v3.11.2 [skip ci]
Fix: Matching N3322 for memcpy UB in C2y
🆕 sz_checksum(char const *, size_t) C 99 interface
sz_checksum(char const *, size_t) C 99 interfacesz::str().checksum() C++ 11 interfacesz.checksum(str) Python interfaceDatabase and other Systems Engineers, you can now use StringZilla to dynamically dispatch different check-sum kernels for AVX2 capable Haswell+ CPUs, AVX-512BW capable Ice Lake+ CPUs, and Arm NEON CPUs on mobile. In AVX-512, masked loads are used extensively, resulting in a 10% improvement even on typical English words, averaging 5 bytes in length and 20x performance improvement compared to the serial code for longer strings.
On the technical side, on x86, the kernels use the well-known SAD(text, zeros) idiom to accumulate absolute differences between individual bytes into 64-bit words. It also uses bidirectional traversal to saturate the core, capable of performing 2 loads per CPU cycle. Moreover, on large inputs, it switches to streaming loads, separately handling the head and the tail, similar to our memcpy alternative, also outperforming LibC on AVX-512-capable machines 😎
sz_checksum visibility (9bec0eb)_mm_cvtsi128_si64x in Clang (c8c6c7c)Fix: Missing C++ function-objects
Release: v3.10.11 [skip ci]
cibuildwheel==2.21.3 (592034c)Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →