NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2200 most downloaded on PyPI
Portable mixed-precision BLAS-like vector math library for x86 and ARM
Last release 7 months ago
07 Mar 2026
Release timing varies
gaps range from 8 days to 6 months
Some releases are documented
notes for 23 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
3 years old
109 releases · first in 2023
One column per quarter.
Nothing published for this version
Most importantly, this is the first SimSIMD release deprecating Python 3.6, released in 2016. Now, 8 years later, we deprecated it to more broadly uti…
Most importantly, this is the first SimSIMD release deprecating Python 3.6, released in 2016. Now, 8 years later, we deprecated it to more broadly utilize the Fast Calling Convention. Read more in a dedicated article on the cost of function arguments parsing in Pyhton - 35% discount on keyword arguments 😄
dtype= issues 👓bf16 dot-product on Arm 🦾[x] Set Intersections & Sparse Distances #100
Major additions:
Minor fixes:
f16, i8, b8 to Python buffers in 33f1b13bf16 to Rust, thanks to @WyctusSome crazy findings:
rsqrt precision on Arm is not documented at allrsqrt approximation for double in AVX-512 is only 6x more accurate, than for floatValueError instead of OverflowErrormath.sqrt over numpy.sqrt when dealing with NumPy arraysqrt in libc is bit-preciseNothing published for this version
This major release adds new capability levels for Arm allowing us to differentiate f16, bf16. and i8-supporting generations of CPUs, becoming increasi
bf16 & Dynamic Dispatch on ArmThis major release adds new capability levels for Arm allowing us to differentiate f16, bf16. and i8-supporting generations of CPUs, becoming increasingly popular in the datacenter. Similar to speedups on AMD Genoa, on Arm Graviton3 the bf16 kernels perform very well:
dot_bf16_neon_1536d/min_time:10.000/threads:1 183 ns 183 ns 76204478 abs_delta=0 bytes=33.5194G/s pairs=5.45563M/s relative_error=0
cos_bf16_neon_1536d/min_time:10.000/threads:1 239 ns 239 ns 58180403 abs_delta=0 bytes=25.7056G/s pairs=4.18386M/s relative_error=0
l2sq_bf16_neon_1536d/min_time:10.000/threads:1 312 ns 312 ns 43724273 abs_delta=0 bytes=19.7064G/s pairs=3.20742M/s relative_error=0
The bf16 kernels reach 33 GB/s as opposed to 19 GB/s for f16:
dot_f16_neon_1536d/min_time:10.000/threads:1 323 ns 323 ns 43311367 abs_delta=82.3015n bytes=19.0324G/s pairs=3.09772M/s relative_error=109.717n
cos_f16_neon_1536d/min_time:10.000/threads:1 367 ns 367 ns 38007895 abs_delta=1.5456m bytes=16.7349G/s pairs=2.72377M/s relative_error=6.19568m
l2sq_f16_neon_1536d/min_time:10.000/threads:1 341 ns 341 ns 41010555 abs_delta=66.7783n bytes=18.0436G/s pairs=2.93679M/s relative_error=133.449n
Arm supports 2x2 matrix multiplications for i8 and bf16. All of our initial attempts with @eknag to use them for faster cosine computations for different length vectors have failed. Old measurements:
cos_i8_neon_16d/min_time:10.000/threads:1 5.41 ns 5.41 ns 1000000000 abs_delta=910.184u bytes=5.91441G/s pairs=184.825M/s relative_error=4.20295m
cos_i8_neon_64d/min_time:10.000/threads:1 7.63 ns 7.63 ns 1000000000 abs_delta=939.825u bytes=16.7729G/s pairs=131.039M/s relative_error=3.82144m
cos_i8_neon_1536d/min_time:10.000/threads:1 101 ns 101 ns 139085845 abs_delta=917.35u bytes=30.394G/s pairs=9.89387M/s relative_error=3.63925m
Attempts with i8 for different dimensionality vectors:
cos_i8_neon_16d/min_time:10.000/threads:1 5.72 ns 5.72 ns 1000000000 abs_delta=0.282084 bytes=5.59562G/s pairs=174.863M/s relative_error=1.15086
cos_i8_neon_64d/min_time:10.000/threads:1 8.40 ns 8.40 ns 1000000000 abs_delta=0.234385 bytes=15.2345G/s pairs=119.02M/s relative_error=0.923009
cos_i8_neon_1536d/min_time:10.000/threads:1 117 ns 117 ns 118998604 abs_delta=0.23264 bytes=26.2707G/s pairs=8.55167M/s relative_error=0.920099
This line of work is to be continued, in parallel with new similarity metrics and distance functions.
f32 spatial distances on older x86 CPUs. Thanks to @makarr 👏.tar.gz packages, that were missing a VERSION file. Thanks to @OD-Ice 👏The "brain-float-16" is a popular machine learning format. It's broadly supported in hardware and is very machine-friendly, but software support is st
bf16 kernelsThe "brain-float-16" is a popular machine learning format. It's broadly supported in hardware and is very machine-friendly, but software support is still lagging behind - https://github.com/numpy/numpy/issues/19808. Most importantly, low-precision bf16 dot-products are supported by the most recent Zen4-based AMD Genoa CPUs. Those have up-to 96 cores, and just one of those cores is capable of computing 86 GB/s worth of such dot-products.
------------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations UserCounters...
------------------------------------------------------------------------------------------------------------
dot_bf16_haswell_1536d/min_time:10.000/threads:1 203 ns 203 ns 68785823 abs_delta=29.879n bytes=30.1978G/s pairs=4.91501M/s relative_error=39.8289n
dot_bf16_haswell_1536b/min_time:10.000/threads:1 93.0 ns 93.0 ns 150582910 abs_delta=24.8365n bytes=33.0344G/s pairs=10.7534M/s relative_error=33.1108n
dot_bf16_genoa_1536d/min_time:10.000/threads:1 71.0 ns 71.0 ns 197340105 abs_delta=23.6042n bytes=86.5917G/s pairs=14.0937M/s relative_error=31.4977n
dot_bf16_genoa_1536b/min_time:10.000/threads:1 36.1 ns 36.1 ns 387637713 abs_delta=22.3063n bytes=85.0019G/s pairs=27.6699M/s relative_error=29.7341n
dot_bf16_serial_1536d/min_time:10.000/threads:1 15992 ns 15991 ns 874491 abs_delta=311.896n bytes=384.216M/s pairs=62.5352k/s relative_error=415.887n
dot_bf16_serial_1536b/min_time:10.000/threads:1 7979 ns 7978 ns 1754703 abs_delta=193.719n bytes=385.045M/s pairs=125.34k/s relative_error=258.429n
dot_bf16c_serial_1536d/min_time:10.000/threads:1 16430 ns 16429 ns 852438 abs_delta=251.692n bytes=373.964M/s pairs=60.8665k/s relative_error=336.065n
dot_bf16c_serial_1536b/min_time:10.000/threads:1 8207 ns 8202 ns 1707289 abs_delta=165.209n bytes=374.54M/s pairs=121.92k/s relative_error=220.35n
vdot_bf16c_serial_1536d/min_time:10.000/threads:1 16489 ns 16488 ns 849194 abs_delta=247.646n bytes=372.639M/s pairs=60.6509k/s relative_error=330.485n
vdot_bf16c_serial_1536b/min_time:10.000/threads:1 8224 ns 8217 ns 1704397 abs_delta=162.036n bytes=373.839M/s pairs=121.693k/s relative_error=216.042n
That's a steep 3x improvement over single-precision FMA throughput we can obtain by simply shifting bf16 left by 16 bits and using _mm256_fmadd_ps intrinsic / vfmadd instruction available since Intel Haswell.
i8 kernelsWe can't directly use _mm512_dpbusd_epi32 every time we want to compute a low-precision integer dot-product, as it's asymmetric with respect to the sign of the input arguments:
Signed(ZeroExtend16(a.byte[4j]) * SignExtend16(b.byte[4j]))
In the past we would just upcast to 16-bit integers and resort to _mm512_dpwssds_epi32. It is a much more costly multiplication circuit, and, assuming that I avoid loop unrolling, also implies 2x fewer scalars per loop. But for cosine distances there is something simple we can do. Assuming that we multiply the vector by itself, even if a certain vector component is negative, its square will always be positive. So we can avoid the expensive 16-bit operation at least where we compute the vector norms:
a_abs_vec = _mm512_abs_epi8(a_vec);
b_abs_vec = _mm512_abs_epi8(b_vec);
a2_i32s_vec = _mm512_dpbusds_epi32(a2_i32s_vec, a_abs_vec, a_abs_vec);
b2_i32s_vec = _mm512_dpbusds_epi32(b2_i32s_vec, b_abs_vec, b_abs_vec);
On Intel Sapphire Rapids it resulted in a higher single-thread utilization, but didn't lead to improvements on other platforms.
---------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations UserCounters...
---------------------------------------------------------------------------------------------------------
cos_i8_haswell_1536d/min_time:10.000/threads:1 92.4 ns 92.4 ns 151487077 abs_delta=105.739u bytes=33.2344G/s pairs=10.8185M/s relative_error=405.868u
cos_i8_haswell_1536b/min_time:10.000/threads:1 92.4 ns 92.4 ns 151478714 abs_delta=0 bytes=33.2383G/s pairs=10.8198M/s relative_error=0
cos_i8_ice_1536d/min_time:10.000/threads:1 61.6 ns 61.6 ns 227445214 abs_delta=0 bytes=49.898G/s pairs=16.2428M/s relative_error=0
cos_i8_ice_1536b/min_time:10.000/threads:1 61.5 ns 61.5 ns 227609621 abs_delta=0 bytes=49.9167G/s pairs=16.2489M/s relative_error=0
cos_i8_serial_1536d/min_time:10.000/threads:1 299 ns 299 ns 46788061 abs_delta=0 bytes=10.2666G/s pairs=3.34198M/s relative_error=0
cos_i8_serial_1536b/min_time:10.000/threads:1 299 ns 299 ns 46787275 abs_delta=0 bytes=10.2663G/s pairs=3.34191M/s relative_error=0
New timings:
---------------------------------------------------------------------------------------------------------
Benchmark Time CPU Iterations UserCounters...
---------------------------------------------------------------------------------------------------------
cos_i8_haswell_1536d/min_time:10.000/threads:1 92.4 ns 92.4 ns 151463294 abs_delta=105.739u bytes=33.2359G/s pairs=10.819M/s relative_error=405.868u
cos_i8_haswell_1536b/min_time:10.000/threads:1 92.4 ns 92.4 ns 151470519 abs_delta=0 bytes=33.2392G/s pairs=10.82M/s relative_error=0
cos_i8_ice_1536d/min_time:10.000/threads:1 48.1 ns 48.1 ns 292087642 abs_delta=0 bytes=63.8408G/s pairs=20.7815M/s relative_error=0
cos_i8_ice_1536b/min_time:10.000/threads:1 48.2 ns 48.2 ns 291716009 abs_delta=0 bytes=63.7662G/s pairs=20.7572M/s relative_error=0
cos_i8_serial_1536d/min_time:10.000/threads:1 299 ns 299 ns 46784120 abs_delta=0 bytes=10.2647G/s pairs=3.34139M/s relative_error=0
cos_i8_serial_1536b/min_time:10.000/threads:1 299 ns 299 ns 46781350 abs_delta=0 bytes=10.2654G/s pairs=3.3416M/s relative_error=0
## 4.3.1 (2024-04-10) ### Make * Lighter Rust package
Dot-products compilation on SVE
## 4.2.1 (2024-03-26) ### Fix * Function attributes (eefb54b) ### Make * Avoid native f16 by default (bd02af2) * Shorter Rust compilation
# 4.2.0 (2024-03-25) ### Add * Compile-time dispatch (b72bf56) ### Fix * Rust benchmarks (5279de7) * Setting signaling NaN (4e6e686) ### Make * Build
## 4.1.1 (2024-03-24) ### Fix * Detecting Ice Lake (098cd90) ### Make * Unused function attributes
This release refactors compiler attributes and intrinsics usage to make it compatible with MSVC. Most noticeably, function defined like this:
This release refactors compiler attributes and intrinsics usage to make it compatible with MSVC. Most noticeably, function defined like this:
__attribute__((target("+simd")))
inline static void simsimd_cos_f32_neon(simsimd_f32_t const* a, simsimd_f32_t const* b, simsimd_size_t n, simsimd_distance_t* result) { }
Now look like this:
#pragma GCC push_options
#pragma GCC target("+simd")
#pragma clang attribute push(__attribute__((target("+simd"))), apply_to = function)
inline static void simsimd_cos_f32_neon(simsimd_f32_t const* a, simsimd_f32_t const* b, simsimd_size_t n, simsimd_distance_t* result) { }
#pragma clang attribute pop
#pragma GCC pop_options
Thanks to that SimSIMD on Windows is gonna be just as fast as on Linux and MacOS 🤗 🪟
vld2_f16 on MSVC (ce9800e)_mm_rsqrt14_ps in MSVC (21e30fe)float16_t in MSVC Arm64 builds (94442c3)This is a packed redesign! Let's start with what's cool about it and later cover the mechanics.
This is a packed redesign! Let's start with what's cool about it and later cover the mechanics.
complex32 Python type ... that:
SimSIMD is now the fastest and most popular library for computing half-precision products/similarities for Fourier Series and other complex data 🥳
What breaks:
AB, instead of 1 - AB for broader applicability.NumPy and SciPy issues with PyPy
crates.io allows only 5 keywords/categories
Don't download Google Benchmark by default
Forward SIMSIMD_NATIVE_F16 (#78) (aff3bcf), closes #78 /github.com/unum-cloud/usearch/blob/ce54b814a8a10f4c0c32fee7aad9451231b63f75/include/usearch/in
Rust cosine function (#77) (cdd282d), closes #77
Nothing published for this version
GoLang bindings (#70) (0627795), closes #70
## 3.6.3 (2024-01-06) ### Make * Revert test location
JS installation, grammar and counters (#50) (ba0e233), closes #50
Update Python library __version__
As was discussed in the SciPy integration thread, Python libraries use double-precision floating-point numbers by default. So in this release I've ext
As was discussed in the SciPy integration thread, Python libraries use double-precision floating-point numbers by default. So in this release I've extended the spatial distance functions - cosine, sqeuclidean, inner with support for double arguments with specialized implementations on AVX-512-capable x86 CPUs and SVE-capable Arm CPUs.
serial, x86_avx2, x86_avx512, x86_avx2fp16, x86_avx512fp16, x86_avx512vpopcntdq, x86_avx512vnniopenblas64dep140640983012528| Datatype | Method | Ops/s | SimSIMD Ops/s | SimSIMD Improvement |
|---|---|---|---|---|
f64 |
scipy.cosine |
63,612 | 572,605 | 9.00 x |
f64 |
scipy.sqeuclidean |
238,547 | 915,596 | 3.84 x |
f64 |
numpy.inner |
449,499 | 986,522 | 2.19 x |
| Datatype | Method | Ops/s | SimSIMD Ops/s | SimSIMD Improvement |
|---|---|---|---|---|
f64 |
scipy.cosine |
68,962 | 1,457,172 | 21.13 x |
f64 |
scipy.sqeuclidean |
247,727 | 1,535,547 | 6.20 x |
f64 |
numpy.inner |
463,509 | 1,512,004 | 3.26 x |
serial, arm_neon, arm_sveopenblas64openblas64| Datatype | Method | Ops/s | SimSIMD Ops/s | SimSIMD Improvement |
|---|---|---|---|---|
f64 |
scipy.cosine |
40,729 | 725,382 | 17.81 x |
f64 |
scipy.sqeuclidean |
160,812 | 728,114 | 4.53 x |
f64 |
numpy.inner |
473,443 | 767,374 | 1.62 x |
f64 |
scipy.jensenshannon |
15,684 | 38,528 | 2.46 x |
f64 |
scipy.kl_div |
49,983 | 61,811 | 1.24 x |
| Datatype | Method | Ops/s | SimSIMD Ops/s | SimSIMD Improvement |
|---|---|---|---|---|
f64 |
scipy.cosine |
41,130 | 1,460,850 | 35.52 x |
f64 |
scipy.sqeuclidean |
162,147 | 1,486,255 | 9.17 x |
f64 |
numpy.inner |
473,856 | 1,580,136 | 3.33 x |
Detecting compile-time capabilities
revert back to using the reciprocal
## 3.5.3 (2023-10-31) ### Make * Upgade JS CI pipeline
Remove any package links for NPM
## 3.5.1 (2023-10-31) ### Make * JS package hard-link resolved
.npmignore & some minor fixes (#37) (f2555af), closes #37
ifnan compilation issues for GBench (fe3286f)f16 to SciPy f64 (dd655c1)Divergence functions are a bit more complex than the Cosine Similarity, primarily because they have to compute logarithms, which are relatively slow w
Divergence functions are a bit more complex than the Cosine Similarity, primarily because they have to compute logarithms, which are relatively slow when using LibC's logf.
So, aside from minor patches, in this PR, I've rewritten the Jensen Shannon distances leveraging several optimizations, mainly focusing on AVX-512 and AVX-512FP16 extensions, which resulted in 4.6x improvement over the auto-vectorized single-precision variant and a whopping 118x improvement over the half-precision code produced by GCC 12.
_mm512_getexp_ph and _mm512_getmant_ph are now used to extract the exponent and the mantissa of the floating-point number, streamlining the process. I've also used Horner's method for the polynomial approximation._mm512_rcp_ph for half-precision and _mm512_rcp14_ps for single-precision. The _mm512_rcp28_ps was found to be unnecessary for this implementation._mm512_cmp_ph_mask is used to compute a mask for close-to-zero values, avoiding the addition of an "epsilon" to every component, which is both cleaner and more accurate._mm512_maskz_fmadd_ph replaces distinct addition and multiplication operations, optimizing the calculation further.To remind, the Jensen Shannon divergence is the symmetric version of the Kullback-Leibler divergence:
JSD(P, Q) = \frac{1}{2} D(P || M) + \frac{1}{2} D(Q || M) \\
M = \frac{1}{2}(P + Q), D(P || Q) = \sum P(i) \cdot \log \left( \frac{P(i)}{Q(i)} \right)
For AVX-512FP16, the current implementation looks like this:
__attribute__((target("avx512f,avx512vl,avx512fp16")))
inline __m512h simsimd_avx512_f16_log2(__m512h x) {
// Extract the exponent and mantissa
__m512h one = _mm512_set1_ph((_Float16)1);
__m512h e = _mm512_getexp_ph(x);
__m512h m = _mm512_getmant_ph(x, _MM_MANT_NORM_1_2, _MM_MANT_SIGN_src);
// Compute the polynomial using Horner's method
__m512h p = _mm512_set1_ph((_Float16)-3.4436006e-2f);
p = _mm512_fmadd_ph(m, p, _mm512_set1_ph((_Float16)3.1821337e-1f));
p = _mm512_fmadd_ph(m, p, _mm512_set1_ph((_Float16)-1.2315303f));
p = _mm512_fmadd_ph(m, p, _mm512_set1_ph((_Float16)2.5988452f));
p = _mm512_fmadd_ph(m, p, _mm512_set1_ph((_Float16)-3.3241990f));
p = _mm512_fmadd_ph(m, p, _mm512_set1_ph((_Float16)3.1157899f));
return _mm512_add_ph(_mm512_mul_ph(p, _mm512_sub_ph(m, one)), e);
}
__attribute__((target("avx512f,avx512vl,avx512fp16")))
inline static simsimd_f32_t simsimd_avx512_f16_js(simsimd_f16_t const* a, simsimd_f16_t const* b, simsimd_size_t n) {
__m512h sum_a_vec = _mm512_set1_ph((_Float16)0);
__m512h sum_b_vec = _mm512_set1_ph((_Float16)0);
__m512h epsilon_vec = _mm512_set1_ph((_Float16)1e-6f);
for (simsimd_size_t i = 0; i < n; i += 32) {
__mmask32 mask = n - i >= 32 ? 0xFFFFFFFF : ((1u << (n - i)) - 1u);
__m512h a_vec = _mm512_castsi512_ph(_mm512_maskz_loadu_epi16(mask, a + i));
__m512h b_vec = _mm512_castsi512_ph(_mm512_maskz_loadu_epi16(mask, b + i));
__m512h m_vec = _mm512_mul_ph(_mm512_add_ph(a_vec, b_vec), _mm512_set1_ph((_Float16)0.5f));
// Avoid division by zero problems from probabilities under zero down the road.
// Masking is a nicer way to do this, than adding the `epsilon` to every component.
__mmask32 nonzero_mask_a = _mm512_cmp_ph_mask(a_vec, epsilon_vec, _CMP_GE_OQ);
__mmask32 nonzero_mask_b = _mm512_cmp_ph_mask(b_vec, epsilon_vec, _CMP_GE_OQ);
__mmask32 nonzero_mask = nonzero_mask_a & nonzero_mask_b & mask;
// Division is an expensive operation. Instead of doing it twice,
// we can approximate the reciprocal of `m` and multiply instead.
__m512h m_recip_approx = _mm512_rcp_ph(m_vec);
__m512h ratio_a_vec = _mm512_mul_ph(a_vec, m_recip_approx);
__m512h ratio_b_vec = _mm512_mul_ph(b_vec, m_recip_approx);
// The natural logarithm is equivalent to `log2`, multiplied by the `loge(2)`
__m512h log_ratio_a_vec = simsimd_avx512_f16_log2(ratio_a_vec);
__m512h log_ratio_b_vec = simsimd_avx512_f16_log2(ratio_b_vec);
// Instead of separate multiplication and addition, invoke the FMA
sum_a_vec = _mm512_maskz_fmadd_ph(nonzero_mask, a_vec, log_ratio_a_vec, sum_a_vec);
sum_b_vec = _mm512_maskz_fmadd_ph(nonzero_mask, b_vec, log_ratio_b_vec, sum_b_vec);
}
simsimd_f32_t log2_normalizer = 0.693147181f;
return _mm512_reduce_add_ph(_mm512_add_ph(sum_a_vec, sum_b_vec)) * 0.5f * log2_normalizer;
}
I conducted benchmarks at both the higher-level Python and lower-level C++ layers, comparing the auto-vectorization on GCC 12 to our new implementation on an Intel Sapphire Rapids CPU on AWS:
<img width="1178" alt="Screenshot 2023-10-23 at 14 33 03" src="https://github.com/ashvardanian/SimSIMD/assets/1983160/39d4b937-b684-4a5f-a010-dafa6dc8d114">
The program was compiled with -O3 and -ffast-math and was running on all cores of the 4-core instance, potentially favoring the non-vectorized solution. When normalized and tabulated, the results are as follows:
| Benchmark | Pairs/s | Gigabytes/s | Absolute Error | Relative Error |
|---|---|---|---|---|
serial_f32_js_1536d |
0.243 M/s | 2.98 G/s | 0 | 0 |
serial_f16_js_1536d |
0.018 M/s | 0.11 G/s | 0.123 | 0.035 |
avx512_f32_js_1536d |
1.127 M/s | 13.84 G/s | 0.001 | 345u |
avx512_f16_js_1536d |
2.139 M/s | 13.14 G/s | 0.070 | 0.020 |
avx2_f16_js_1536d |
0.547 M/s | 3.36 G/s | 0.011 | 0.003 |
Of course, the results will vary depending on the vector size. I generally use 1536 dimensions, matching the size of OpenAI Ada embeddings, standard in NLP workloads. The Jensen Shannon divergence, however, is used broadly in other domains of statistics, bio-informatics, and chem-informatics, so I'm adding it as a new out-of-the-box supported metric into USearch today 🥳
This further accelerates the k-approximate Nearest Neighbors Search and the clustering of Billions of different protein sequences without alignment procedures. Expect one more "Less Slow" post soon! 🤗
Nothing published for this version
Nothing published for this version
functions for probability distributions
## 2.3.8 (2023-10-10) ### Improve * Use fast calling convention
Type-conversion in default distance functions
Compile time __version__ definition
## 2.1.2 (2023-10-08) ### Make * Use default compiler on MacOS
Division by zero in cosine distance
bench.py and expose Hamming and Jaccard
bench.py and expose Hamming and Jaccard (8fad89f)f32 (4d40999)_Float16 compilation issue (3847b30)_Float16 support on x86 + GCC (10517ee)avx512vpopcntdq capability (8724f9e)sqrt include (630234c)avx512_f32_ip (7ff54ca)simsimd_sve_f16_l2sq (02c6cfd)## 2.0.4 (2023-10-05) ### Make * Avoid uncommon Python version
Disable python universal builds
Placeholders for string similarity
avx512bw label for masked loads (ccce7c8)avx512f flag for __m512 registers (14ceef0)avx512fp16 compilation flag (f8640e0)fma flag for AVX2 cosine and dot-product (850a0e8)_mm256_popcnt_epi64 (9055a02)union instantiations (ebc4b12)Your coding agent can read these notes before it upgrades. Set up the MCP server →