NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3097 most downloaded on PyPI
AI coding assistant skill (Claude Code, CodeBuddy, Codex, OpenCode, Kilo Code, Cursor, Gemini CLI, Aider, OpenClaw, Factory Droid, Trae, Hermes, Kiro, Pi, Devin CLI, Google Antigravity) - turn any folder of code, docs, papers, images, or videos into a queryable knowledge graph
Last release today
04 Oct 2026
Ships on a steady schedule
a new release about every 8 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
6 months old
244 releases · first in 2026
One column per month.
Fix: LLM empty choices / None message guard — Gemini and other providers return choices=[] on content-filtered HTTP 200 responses; now raises a clean
choices=[] on content-filtered HTTP 200 responses; now raises a clean error instead of crashing with IndexError (#924)general-purpose agent reference and headless-incompatible interactive halt (#911, closes #825)Fix: git hooks phantom directory on git < 2.31
--resolution N for extract and cluster-only (#919)--exclude-hubs P for extract and cluster-only with majority-vote reattachment (#919)--path-format=absolute, validate path contains no newlines, anchor relative paths on repo root (#907)save_manifest incremental data loss — seed from existing manifest before loop so untouched files aren't erased on partial runs (#917)base_class_clause for class_specifier and struct_specifier (#915)cohesion_score now returns raw float, display rounds to 2dp (#919)Type::method() scoped calls and common trait-method names from cross-file resolver (#908)--resolution N for extract and cluster-only — control Leiden/Louvain community granularity (>1 = more smaller, <1 = fewer larger) (#919)--exclude-hubs P for extract and cluster-only — exclude degree-percentile super-hubs from partitioning, reattach by majority-vote neighbour community (#919)Feat: DeepSeek backend support — set DEEPSEEK_API_KEY and use --backend deepseek; default model deepseek-v4-flash
DEEPSEEK_API_KEY and use --backend deepseek; default model deepseek-v4-flashShipping a new subcommand today: graphify prs.
Shipping a new subcommand today: graphify prs.
The problem it solves: no existing PR dashboard knows your codebase structure. You can see CI state and review decisions — but you cannot see that two open PRs both touch the auth community and are going to conflict at merge time. graphify prs does.
graphify prs # dashboard: CI state, review decision, worktree mapping
graphify prs 42 # deep dive on one PR — blast radius, communities touched
graphify prs --conflicts # PRs sharing graph communities → merge-order risk
graphify prs --triage # AI ranks your review queue by graph impact
graphify prs --worktrees # worktree → branch → PR in one view
graphify prs --repo owner/repo # works on any GitHub repo
The --conflicts view crosses your open PRs against the knowledge graph. If two PRs touch the same community of nodes, they get flagged — with representative node labels so you know what that community actually is (ValidateToken, SessionStore, AuthMiddleware) rather than just "Community 3".
The dashboard shows a blast radius column per PR: 14 nodes / 3 communities. High number = review this carefully before merging.
--triage sends your PR queue to whatever LLM backend you have configured and gets a ranked list with one action per PR. Auto-detects from your env: claude → kimi → openai → gemini → claude-cli → ollama. Override with GRAPHIFY_TRIAGE_BACKEND.
Three new tools on the MCP server: list_prs, get_pr_impact, triage_prs. Your agent can now query PR state the same way it queries the graph.
pip install --upgrade graphifyy
No new required dependencies. Requires gh CLI authenticated for PR data.
graphify prs — graph-aware PR dashboard: CI state, review decision, worktree mapping, and graph blast radius per PR; --triage ranks your queue via any configured LLM backend (claude, kimi, openai, gemini, claude-cli, ollama — auto-detected); --conflicts shows PRs sharing graph communities with node labels; --worktrees maps worktree paths to branches to open PRs; MCP tools list_prs, get_pr_impact, triage_prs for agent accessIDF weighting — common terms like error or handle that match dozens of nodes are down-weighted, so a rare identifier like FooBarService ranks first an
error or handle that match dozens of nodes are down-weighted, so a rare identifier like FooBarService ranks first and BFS expands from the right starting pointget_node or add a context_filter) rather than just reporting truncationint x;, static const int MAX = 100;) are now extracted as nodes with defines edges from the parent class — the previous field_declaration branch was silently a no-op due to a wrong child type filter_get_cpp_func_name now handles field_identifier, destructor_name, and operator_name node types#include "path/to/file.h" edges now resolve relative to the including file and use the full resolved path as the target node ID — matching what extraction creates for the included file. Previously all include edges dangled with a basename-only ID and were silently dropped.handle, init, run) from different files no longer collapse into artificial god nodes; cross-file matches fall through to Pass 2 fuzzyexact_merges counter now reports only merges actually performeduvx graphifyy@0.8.7 install
# or
pip install --upgrade graphifyy
error or handle that match dozens of nodes are down-weighted so a rare identifier like FooBarService ranks first and BFS expands from the right node (#897)query_graph now tells Claude what to do (call get_node or add a context_filter) rather than just saying "truncated" (#897)int x;, static const int MAX = 100;) now extracted as nodes with defines edges from the parent class — previously the field_declaration branch was a no-op due to a wrong child type guard (#898)handle, init, run) from different files no longer collapse into artificial god nodes; cross-file matches are routed to Pass 2 fuzzy (#895)#include "path/to/file.h" edges now resolve the include path relative to the including file and use the full resolved path as the target node ID, matching what extraction creates for the included file — previously all include edges dangled with a basename-only ID (#899)exact_merges counter in dedup now reports only merges actually performed rather than counting all same-label nodes across files (#895)Two new suppression rules stop false positives from dominating the headline output:
Two new suppression rules stop false positives from dominating the headline output:
AuthError → TypeScript Member). These calls/uses edges are now zero-scored so they don't crowd out real structural surprises. semantically_similar_to and EXTRACTED edges are unaffected.calls edge, that's documentation cross-reference noise, not architecture. Same suppression applied. Code↔paper edges are preserved (a code file referencing a research paper is a genuine cross-format signal).name, id, type, start, end, key, value, data, items, title, description, version, properties) accumulate positional degree from sibling records rather than architectural meaning. These are now excluded from god_nodes. Domain-specific labels in JSON files still rank normally.upgrade()) are still captured.follow_symlinks is now enabled automatically when symlinked children are detected in the target directory, no flag needed./graphify query interactively rather than read GRAPH_REPORT.md first.calls/uses edges (e.g. Python → TypeScript) are suppressed in Surprising Connections — label-matching across language boundaries in monorepos is resolver pollution, not structural insight; all structural bonuses zeroed for these edgescalls/uses edges suppressed in Surprising Connections — the LLM seeing a symbol name in a README and emitting a calls edge is documentation cross-reference noise, not a real architectural connection (#890)name, id, type, start, end, key, value, data, items, title, description, version, properties) filtered from god_nodes — their degree is positional (every sibling record in the same JSON file references them), not architectural (#890)--follow-symlinks is now auto-detected — if symlinked children are present in the target directory, follow-symlinks is enabled automatically without requiring an explicit flag (#887)/graphify query interactively rather than reading GRAPH_REPORT.md first; the report is a summary, not a starting point (#891)Gitignore parent-exclusion rule (#882): .graphifyignore patterns now correctly exclude files under an excluded directory even when a ! negation exists
.graphifyignore patterns now correctly exclude files under an excluded directory even when a ! negation exists elsewhere in the file. Previously, any negation pattern would disable directory pruning entirely.ASR1603/ASR1605 or M1/M1 Pro. Two new guards (_is_variant_pair, _short_label_blocked) prevent these false positives while still catching real typos.--update now correctly removes nodes and edges from deleted source files (was matching on basenames instead of full paths).graph.json are now coerced to int before community label lookup, fixing blank community names in reports and HTML.--update with deletions-only no longer errors on missing extraction file.uv tool and pipx installs on Windows (#831)..graphifyignore parent-exclusion rule now correctly blocks files under an excluded directory even when a ! negation exists elsewhere in the file — previously any negation pattern disabled directory pruning entirely (#882)ASR1603/ASR1605 or M1/M1 Pro — Jaro-Winkler prefix bonus is now gated by _is_variant_pair and _short_label_blocked guards; real typos on short labels still merge (#878)worked/rsl-siege-manager/ — case study on a real-world Python + TypeScript monorepo (FastAPI backend, React/Vite frontend, Discord bot); covers god node behaviour with tests included, cross-language INFERRED edges, community cohesion, and Alembic migration noise (#881)Feat: Firebird SQL — trigger and stored procedure extraction via CREATE TRIGGER and regex fallback; FK detection via global regex covering REFERENCES
CREATE TRIGGER and regex fallback; FK detection via global regex covering REFERENCES and FOREIGN KEY clauses (#875)--update deletion pruning now matches on full source file paths instead of basenames, preventing false node removal when different directories contain files with the same name (#876)--update now also prunes edges whose source_file attr points to deleted files, not just nodes (#876)graph.json (stored as strings) are now coerced to int before lookup, fixing blank community names in GRAPH_REPORT.md and graph.html (#877)Fix: Windows skill temp files (chunk JSONs, .graphify_python, .graphify_root) no longer pollute the project root — all written under graphify-out/
.graphify_python, .graphify_root) no longer pollute the project root — all written under graphify-out/ (#831)--update with deletions-only no longer errors when .graphify_extract.json does not yet exist — creates an empty extraction file before merging (#876)Fix: Python interpreter detection for uv tool and pipx installs on Windows — graphify install and all skill steps now find the correct executable
uv tool and pipx installs on Windows — graphify install and all skill steps now find the correct executable (#831).github/, .vscode/) are now indexed when explicitly included via .graphifyignore (#873)graph.json changes on disk (#874)Shell scripts and JSON configs are now indexed automatically via tree-sitter — no LLM, no tokens, no flags needed.
Shell scripts and JSON configs are now indexed automatically via tree-sitter — no LLM, no tokens, no flags needed.
Bash extracts functions, cross-function calls, source imports resolved to real file paths, and export/declare variables.
JSON extracts the full key tree, dependencies blocks as import edges, extends chains (tsconfig, eslintrc), and $ref references.
graphify export callflow-html
Generates a self-contained HTML page with Mermaid call-flow diagrams per module. With graphify hook install set up, the diagram regenerates automatically on every git commit.
.graphifyignore (#861)graphify update no longer forces a full re-extract after an AST-only run (#857)pip install --upgrade graphifyy
Full changelog: https://github.com/safishamsi/graphify/blob/v8/CHANGELOG.md
.sh and .bash files now indexed via tree-sitter; extracts functions, cross-function calls, source/. imports resolved to real file paths, and export/declare variable declarations (#866).json files now indexed via tree-sitter; extracts key/value contains tree, dependencies/devDependencies blocks as imports edges, extends edges (tsconfig, eslintrc), and $ref references (#866).sh, .bash, .json added to CODE_EXTENSIONS in detect.py so files are picked up during corpus scan (#866)*-callflow.html exists in graphify-out/ — works with --watch and graphify hook installcoverage/, lcov-report/, visual-tests/, visual-test/, __snapshots__/, snapshots/, storybook-static/, dist-protected/ added to _SKIP_DIRS — generated artefact dirs no longer appear in the corpus (#869, #870)graphify hook install now works in git linked worktrees — uses git rev-parse --git-path hooks instead of constructing .git/hooks/ directly (#865)graphify-out/converted/ are now checked against .graphifyignore before being added to the file list (#861)save_manifest() accepts a kind parameter (ast, semantic, both) — incremental AST-only graphify update no longer overwrites semantic_hash entries, preventing spurious full re-extracts on the next run (#857)skill-windows.md Step B3 were missing the graphify-out/ prefix, causing chunk files to be written to the wrong directory (#862)`.astro` support — frontmatter static imports, dynamic imports, and \ \ block imports all produce edges; tsconfig path aliases resolved
.astro support — frontmatter static imports, dynamic imports, and `<script>` block imports all produce edges; tsconfig path aliases resolved (#850).rebuild.lock orphan — lock file now contains a single PID while running and is unlinked on release; no more PID accumulation or orphan locks blocking downstream tooling (#858)``` pip install -U graphifyy
uv tool install graphifyy ```
.astro files now extracted as code — frontmatter static imports, dynamic imports, and <script> block imports all produce edges; tsconfig path aliases resolved (#850, PR #852).rebuild.lock no longer accumulates PIDs across rebuilds — now contains a single owning PID while running and is unlinked on release so downstream tooling polling for its absence unblocks promptly (#858, PR #859)ANTHROPIC_API_KEY or other provider keys during /graphify skill runs — the host IDE session provides the LLM (PR #864)`graphify update` is now idempotent — graph.json and GRAPH_REPORT.md only rewritten when content actually changes; topology comparison short-circuits
graphify update is now idempotent — graph.json and GRAPH_REPORT.md only rewritten when content actually changes; topology comparison short-circuits clustering on unchanged graphs (#824)graphify update --no-cluster — skip reclustering, write raw AST graph only (mirrors graphify extract --no-cluster) (#824)--no-cluster schema — now writes "links" key consistent with the full clustered path; previously toggled to "edges" on every mode switch.graphify_labels.json no longer rewritten on no-op rebuilds_check_shrink() helper across both code paths{parent_dir}_{filename_stem}_{entity}; the old filename-only format caused ghost-duplicate nodes when AST and semantic extractors disagreed; existing graphs with ghost duplicates can be cleaned with graphify extract --forcedefault=str in clustering sort keys prevents crashes on non-JSON-serializable edge attributespip install -U graphifyy
# or
uv tool install graphifyy
graphify update is now idempotent — graph.json and GRAPH_REPORT.md are only rewritten when content actually changes; topology comparison short-circuits clustering entirely on unchanged graphs, eliminating residual community-count drift (#824).graphify_labels.json labels don't drift onto wrong communities (#824)--no-cluster flag added to graphify update — writes raw AST graph without clustering, consistent with graphify extract --no-cluster (#824)graphify update --no-cluster now writes "links" key matching the schema of the full clustered path; previously wrote "edges", causing schema toggle on every mode switch.graphify_labels.json was rewritten on every rebuild even when nothing changed; now only written when outputs actually change_check_shrink() helper{parent_dir}_{filename_stem}_{entity} — the old filename-only format caused ghost-duplicate nodes when AST and semantic extractors disagreed on the stem; top-level files use just the filename stem; existing graphs with ghost duplicates can be cleaned up with graphify extract --forcedefault=str) prevents crashes when edge attributes contain non-serializable values`--backend claude-cli` — routes extraction through the Claude Code CLI, no API key needed, zero cost (#855/#856)
--backend claude-cli — routes extraction through the Claude Code CLI, no API key needed, zero cost (#855/#856)--> and <-- now reflect actual edge direction in directed graphs (#849/#853)shortest_path and get_neighbors MCP tools (#849/#853)--update manifest shrink — incremental re-runs no longer drop previously-indexed files (#837)file_type enum — prompt, schema, and synonym mapper all use the same 6 values (#840)ANTHROPIC_API_KEY (#846)sample.F90 → sample_preprocessed.F90 (#823)pip install -U graphifyy
# or
uv tool install graphifyy
graphify path and graphify explain now render arrow direction correctly — --> for caller→callee, <-- for callee←caller; previously the graph was loaded undirected so every hop printed --> regardless of stored direction (#849, #853)shortest_path and get_neighbors tools had the same reversed-arrow bug; now fixed in serve.py alongside the CLI commands (#849, #853)graphify extract --backend bedrock was rejected by the CLI guard even when AWS_PROFILE/AWS_REGION/AWS_DEFAULT_REGION/AWS_ACCESS_KEY_ID were set — boto3 session auth was never reached (#846)max(50, p99_degree)) as transit — hubs can still be destinations but no longer produce semantically meaningless 2-hop paths like ClassA → View → ClassB in Android/Spring corpora (#830)--update manifest shrink — after an incremental run, manifest.json was overwritten with only the changed-file subset, causing the next --update to re-flag the entire unchanged corpus as new; Step 9 now persists the full corpus via all_files fallback (#837)file_type enum aligned across skill.md and llm.py (both now enumerate all six values: code, document, paper, image, rationale, concept); synonym mapper in build.py silently coerces known LLM-emitted synonyms (pattern→concept, markdown→document, tool→code, etc.) before validation (#840)sample.F90 → sample_preprocessed.F90 to avoid case-collision with sample.f90 on macOS case-insensitive filesystems (credit: @FatahChan, #823)`bash pip install -U graphifyy `
read_text()/write_text() calls in the skill pipeline now specify encoding="utf-8" — bare calls defaulted to system codepage on Chinese-locale Windows, silently mojibaking node labels and Markdown content on --update. json.dumps now uses ensure_ascii=False so CJK characters are stored as-is.uv tool install --upgrade graphifyy over pip when uv is on PATH — pip was installing to the wrong environment when graphify was originally installed via uv tool._score_nodes now uses three-tier precedence (exact 1000 / prefix 100 / substring 1) — graphify path "Foo" "FooBar" no longer returns 0 hops when both labels substring-match the same node. Clear error emitted when source and target resolve to the same node.file_hash normalises path keys via .as_posix().lower() — Windows junction/case variants hash identically, fixing save_semantic_cache always reporting "Cached 0 files" on subsequent --update runs. check_semantic_cache now mirrors the same abs-path normalization as the save side._AGENTS_MD_SECTION now includes the /graphify skill trigger instruction — all 7 AGENTS.md platforms (OpenCode, Codex, Aider, Trae, Hermes, OpenClaw, Factory Droid) now correctly invoke the skill tool when the user types /graphify.pip install -U graphifyy
read_text()/write_text() calls in skill.md and skill-windows.md now specify encoding="utf-8" — bare calls defaulted to the system codepage on Chinese-locale Windows, silently mojibaking node labels and Markdown content on --update (#832)json.dumps in skill pipeline now uses ensure_ascii=False so Chinese/CJK characters are stored as-is rather than \uXXXX escaped (#832)uv tool install --upgrade graphifyy over pip when uv is on PATH — pip was installing to the wrong environment when graphify was originally installed via uv tool (#831)_score_nodes in serve.py now uses three-tier precedence (exact 1000 / prefix 100 / substring 1) instead of flat substring scoring — graphify path "Foo" "FooBar" no longer returns 0 hops when both labels substring-match the same node (#828)graphify path and MCP _tool_shortest_path now emit a clear error when source and target resolve to the same node, instead of silently returning 0 hops (#828)file_hash in cache.py now normalises path keys via .as_posix().lower() — Windows junction/case variants of the same file now hash identically, fixing save_semantic_cache always reporting "Cached 0 files" on subsequent --update runs (#826)check_semantic_cache now applies the same absolute-path normalization as save_semantic_cache so relative source_file paths resolve consistently on both sides (#826)_AGENTS_MD_SECTION now includes the /graphify skill trigger instruction — all 7 AGENTS.md platforms (OpenCode, Codex, Aider, Trae, Hermes, OpenClaw, Factory Droid) now correctly invoke the skill tool when the user types /graphify (#827)`bash pip install -U graphifyy `
-h/--help/-? guard: any help flag in any position now stops execution and prints usage — previously graphify cursor install --help silently installed into Cursor, and graphify benchmark --help crashed with FileNotFoundError: '--help'--version, -v, and graphify version now print the installed version and exit — previously fell through to "unknown command"GRAPHIFY_OLLAMA_NUM_CTX=<invalid> no longer falls back to hardcoded 131072 (which exhausted VRAM on constrained cards) — now falls through to auto-derived value with a warningGRAPHIFY_OLLAMA_NUM_CTX is pinned smaller than the estimated chunk size, graphify now warns explicitly that Ollama will silently truncate the prompt and suggests a corrected --token-budgetpip install -U graphifyy
-h/--help/-? in any position now stops execution — previously graphify cursor install --help silently installed into Cursor; graphify benchmark --help crashed with FileNotFoundError (#821)--version, -v, and graphify version now print the installed version and exit (#818)GRAPHIFY_OLLAMA_NUM_CTX=<invalid> no longer falls back to hardcoded 131072 (which exhausted VRAM) — it now falls through to the auto-derived value and prints a warning (#820)GRAPHIFY_OLLAMA_NUM_CTX is set smaller than the estimated chunk size, graphify now warns explicitly that Ollama will silently truncate the prompt and suggests a corrected --token-budget (#820)Skill version mismatch warning suppressed during hook-check (runs on every editor tool use, must be silent) and routed to stderr for all other command
_make_id now uses [^\w]+ with re.UNICODE to preserve non-ASCII word chars. Added NFKC normalization so composed/decomposed forms of the same character produce the same ID. Switched to casefold() for correct locale-sensitive lowercasing. _normalize_id in build.py kept in sync.or so empty-string source isn't silently swapped for the from fallback. Stale from/to keys are now popped after remapping so they can't leak into edge attributes in graph.json.--update merge block in skill.md now calls build_merge() directly instead of an inline NetworkX round-trip that re-introduced the direction-flip bug from #760. Dict merge ordering fixed so explicit source/target always win over stale attrs. Hyperedges pulled from G.graph (full merged set) rather than just the new extraction.CHUNK_PATH derived from graphify-out/.graphify_root at dispatch time — chunk files no longer silently land in the wrong directory due to undefined subagent cwd.hook-check (runs on every editor tool use, must be silent) and routed to stderr for all other commands.pip install -U graphifyy
_make_id and _normalize_id now apply NFKC Unicode normalization before ID generation -- composed/decomposed forms of the same character (e.g. é typed vs pasted from a PDF) now produce the same node ID; switched from .lower() to .casefold() for correct Turkish/German/Greek case folding; both functions are now byte-for-byte equivalent (#811)[^\w]+ with re.UNICODE replaces the old [^a-zA-Z0-9]+ so Unicode word chars are preserved as part of the ID (#811)or so empty-string source is not silently swapped for from; stale from/to keys are now popped before the edge is emitted so they can't leak into graph.json edge attributes (#803)--update merge now calls build_merge() directly instead of an inline NetworkX round-trip that re-introduced the direction-flip bug from #760; dict merge ordering fixed so explicit source/target always win over stale attrs; hyperedges pulled from G.graph (merged) rather than just the new extraction (#801)CHUNK_PATH injected at dispatch time from graphify-out/.graphify_root) so the Write tool doesn't lose chunks to an undefined working directory (#808)hook-check (runs on every editor tool use and must be silent) and routed to stderr for all other commandsOllama VRAM exhaustion (#798): num_ctx is now derived from the actual chunk size instead of hardcoded 131072. With --token-budget 8192, the old value
num_ctx is now derived from the actual chunk size instead of hardcoded 131072. With --token-budget 8192, the old value forced Ollama to allocate 128k KV-cache slots on a 31B model — 4×128k slots by chunk 4 caused OOM. New formula: min(input_tokens + output_cap + 2000, 131072) so an 8k chunk gets ~26k instead.GRAPHIFY_OLLAMA_NUM_CTX / GRAPHIFY_OLLAMA_KEEP_ALIVE env vars as tuning knobs.graphify export callflow-html (#797): generates a self-contained Mermaid architecture/call-flow HTML page from graphify-out/graph.json — community sections, interactive flowcharts with zoom/pan, call detail tables, and graph report highlights.--watch rebuild and post-commit hook if the file already exists. Run once, stays current forever.uv tool upgrade graphifyy
# or: pip install --upgrade graphifyy
num_ctx now derived from actual chunk size instead of hardcoded 131072 -- over-allocating 128k KV-cache slots for small chunks exhausted VRAM by chunk 4 on large models; formula is min(input_tokens + output_cap + 2000, 131072) so --token-budget 8192 gets ~26k instead of 131072 (#798)GRAPHIFY_OLLAMA_NUM_CTX / GRAPHIFY_OLLAMA_KEEP_ALIVE env vars as tuning knobs (#798)graphify export callflow-html -- generates a self-contained Mermaid architecture/call-flow HTML page from graphify-out/graph.json, grouped by community with interactive zoom/pan diagrams, call detail tables, and graph report highlights (#797)--watch rebuild and post-commit hook if the file already exists -- opt-in by existence, zero config (#800)Fix: context-window retry — API calls that fail with context_length_exceeded now automatically bisect the file chunk and retry, up to 6 levels deep. N
context_length_exceeded now automatically bisect the file chunk and retry, up to 6 levels deep. No more hard failures on large files. (#789)print_benchmark() falls back to ASCII on cp1252 consoles; BrokenProcessPool (spawn mode without __main__ guard) now falls back to sequential extraction instead of crashing; Windows skill file rewrites all python -c "..." blocks as PowerShell heredocs to fix quote-escaping failures. (#788)calls edges after --update — build_merge() now reads saved JSON directly instead of round-tripping through NetworkX, which was silently reversing edge direction on reload. (#760)os.replace() prevents half-installed empty skill directories; version-stamp guard and warning added for missing installs. (#725)graphify uninstall — removes skill files from all platforms in one shot; --purge also deletes graphify-out/.ALTER TABLE FK extraction — ADD CONSTRAINT ... FOREIGN KEY and ADD FOREIGN KEY DDL now emit references edges; schema-qualified table names resolved correctly. (#779)uv tool upgrade graphifyy
# or
pip install --upgrade graphifyy
"context_length_exceeded", "maximum context length", and "too_large" across OpenAI-compat backends (#789)print_benchmark() falls back to ASCII box-drawing on cp1252 consoles; ProcessPoolExecutor BrokenProcessPool caught and falls back to sequential extraction when caller lacks if __name__ == "__main__": guard; Windows skill file (skill-windows.md) rewrites all python -c "..." blocks as PowerShell heredocs to fix quote-escaping failures (#788)calls edges after --update -- build_merge() now reads the saved JSON directly instead of round-tripping through NetworkX node_link_graph(), which was silently reversing edge direction on reload (#760)os.replace() pattern prevents half-installed empty skill directories that looked valid but contained no file; version-stamp guard and warning added for missing installs (#725)graphify uninstall top-level command -- removes graphify skill files from all platforms in one shot; --purge flag also deletes graphify-out/ALTER TABLE FK extraction -- ADD CONSTRAINT ... FOREIGN KEY and ADD FOREIGN KEY DDL statements now emit references edges; schema-qualified table names (schema.table) correctly resolved (#779)Fix: .tsx files now use language_tsx grammar for JSX-aware parsing -- previously language_typescript was used, silently dropping all JSX-specific node
.tsx files now use language_tsx grammar for JSX-aware parsing -- previously language_typescript was used, silently dropping all JSX-specific nodes from .tsx files (#766)edges key in saved graph JSON now normalised to links before loading -- prevents KeyError: 'links' on graphs written by older NetworkX versions in query, path, explain, and serve (#768)gws export drops unsupported resourceKey query param -- Drive API requires it as an HTTP header; sending it as a query param was a silent no-op (#772)\n/\r; YAML frontmatter escapes U+2028, U+2029, tabs, and C0; MCP sanitize_label applied to all LLM-derived fields; C preprocessor blocked from #include exfiltration via -nostdinc -I /dev/null; merge-driver 50 MB file size cap and 100k node cap; detect_backend() places Ollama last so paid API keys take precedence over ambient OLLAMA_BASE_URL; Neo4j --password reads from NEO4J_PASSWORD env var by default; hooks exception handling narrowed to (configparser.Error, OSError)CLAUDE.md / AGENTS.md / GEMINI.md templates strengthened with ALWAYS/NEVER/IF ... EXISTS graph-first directives (#775)pip install -U graphifyy
Feat: TypeScript extraction parity -- interface, enum, type alias, and module-level const nodes extracted; new_expression emits calls edges; parity wi
.qmd) file support -- routed through existing Markdown extractor (#761)graphify extract ./docs --google-workspace converts .gdoc, .gsheet, and .gslides files into Markdown sidecars via the gws CLI; account email pseudonymized via SHA256 hash; [google] extra adds Sheets table rendering support (#752)graphify extract ./docs --backend bedrock; credentials via standard AWS provider chain (AWS_PROFILE, AWS_REGION, IAM roles, SSO); model via GRAPHIFY_BEDROCK_MODEL (default anthropic.claude-3-5-sonnet-20241022-v2:0); [bedrock] extra adds boto3 (#757)pip install -U graphifyy
.qmd) file support -- routed through existing Markdown extractor; Quarto executable code blocks (```{python}) extracted as code nodes (#761)graphify extract ./docs --google-workspace converts .gdoc, .gsheet, and .gslides files into Markdown sidecars with the gws CLI before semantic extraction; account email pseudonymized via SHA256 hash; [google] extra adds Sheets table rendering support (#752)gws from the sidecar output directory with a relative -o path, matching gws path validation and avoiding failures when extracting a corpus outside the current working directory.graphify extract ./docs --backend bedrock; credentials via standard AWS provider chain (AWS_PROFILE, AWS_REGION, IAM roles, SSO); model via GRAPHIFY_BEDROCK_MODEL (default anthropic.claude-3-5-sonnet-20241022-v2:0); [bedrock] extra adds boto3 (#757)Gemini + OpenAI backends — graphify extract ./docs --backend gemini (GEMINI_API_KEY / GOOGLE_API_KEY) or --backend openai (OPENAI_API_KEY); pip instal
graphify extract ./docs --backend gemini (GEMINI_API_KEY / GOOGLE_API_KEY) or --backend openai (OPENAI_API_KEY); pip install graphifyy[gemini] / graphifyy[openai].groovy and .gradle via tree-sitter-groovy; Spock def "feature"() syntax handled via regex fallback.luau (Roblox Luau) extracted using the Lua parser.md/.mdx files (zero new deps)collect_files() extension sync — 18 extensions (.sql, .vue, .svelte, .jsx, .ex, .jl, etc.) were silently skipped in skill-mode extraction; now auto-syncs with _DISPATCH.svelte.ts, .svelte.js, index.ts directory, and multi-dot imports now resolve correctlycluster-only now loads and saves .graphify_labels.json — human labels survive re-clustering (#744)graphify export wiki fails fast (exit 1) when .graphify_analysis.json is missing — prevents silent wiki deletion (#746)detect_incremental forwards follow_symlinks — symlinked subtrees no longer vanish on --update (#736)pip install openai (#750)uv tool upgrade graphifyy
pip install --upgrade graphifyy
require() imports now extracted from JS/TS -- const { foo } = require('./mod'), const m = require('./mod'), and const x = require('./mod').y all emit EXTRACTED imports_from (and per-symbol imports) edges. Previously CJS-only Node.js codebases produced AST graphs missing every import edge, which downgraded all cross-file calls to INFERRED.calls edges are now promoted from INFERRED to EXTRACTED when the caller's file has an explicit imports or imports_from edge to the callee. Previously every cross-file call was unconditionally INFERRED, even when a top-of-file import / require proved the binding. On a 92-file CJS Node.js corpus this promoted 88% of cross-file calls (104 of 118) to EXTRACTED.graphify extract ./docs --backend gemini (GEMINI_API_KEY / GOOGLE_API_KEY) or --backend openai (OPENAI_API_KEY); [gemini] and [openai] extras added (#735).groovy and .gradle extracted via tree-sitter-groovy; Spock spec files (def "feature"() syntax) handled via regex fallback (#732).luau (Roblox Luau) added to code extraction using the Lua tree-sitter parser (#745).md and .mdx files with zero new dependencies (#711)collect_files() extension set now auto-syncs with _DISPATCH -- previously 18 extensions (.sql, .vue, .svelte, .jsx, .ex, .jl, etc.) were silently skipped in skill-mode extraction (#711)detect_incremental now forwards follow_symlinks to detect() -- symlinked subtrees no longer vanish on --update runs (#736).svelte.ts / .svelte.js / index.ts directory / multi-dot imports now resolve correctly -- previously these produced phantom edges dropped at merge time (#717, #716)cluster-only now loads and saves .graphify_labels.json -- human-readable community labels survive re-clustering instead of resetting to "Community N" (#744)graphify export wiki now fails fast with exit 1 if .graphify_analysis.json is missing -- prevents silent deletion of existing wiki articles (#746)to_wiki() now raises before the cleanup loop when communities is empty -- second safety layer against wiki data loss (#746)pip install openai; [ollama] extras group added (#750)Ollama backend — graphify extract ./docs --backend ollama; auto-detected when OLLAMA_BASE_URL is set; defaults to qwen2.5-coder:7b; zero cost; pip ins
graphify extract ./docs --backend ollama; auto-detected when OLLAMA_BASE_URL is set; defaults to qwen2.5-coder:7b; zero cost; pip install graphifyy[ollama]graphify global add/remove/list/path registers multiple project graphs at ~/.graphify/global.json with <repo>::<id> prefixed node IDs preventing silent collisions--global --as <tag> flag on graphify extract — registers into the global graph in one stepmerge-graphs collision safety — inputs are now prefix-relabeled before composing, preventing silent node ID collisionsdeduplicate_entities raises ValueError if called with nodes spanning multiple repos (cross-project dedup disabled by design)--dedup-llm flag now correctly threads through to both fresh and incremental extract pathsgraphify antigravity install writes to .agent/ (not .agents/)merge-graphs now accepts both edges and links keys in graph.jsonuv tool upgrade graphifyy
pip install --upgrade graphifyy
graphify extract ./docs --backend ollama; auto-detected when OLLAMA_BASE_URL is set; defaults to qwen2.5-coder:7b; zero cost ($0.00); sentinel API key handles OpenAI client auth requirement (#729)~/.graphify/global.json -- graphify global add/remove/list/path to register multiple project graphs with <repo>::<id> prefixed node IDs, preventing silent collisions; hash-based skip avoids re-ingesting unchanged graphs (#729)graphify extract --global --as <tag> flag -- after building a project graph, auto-registers it into the global graph in one step (#729)merge-graphs now prefix-relabels each input graph before composing, preventing silent node ID collisions when two projects share entity names (#729)deduplicate_entities raises ValueError if called with nodes spanning multiple repos (cross-project dedup disabled by design -- per-project graphs are deduplicated in isolation) (#729)detect_incremental() now accepts and forwards follow_symlinks to detect(). Without this, --update runs silently miss any files reached through a symlinked sub-tree (e.g. state_of_truth/ symlinking to a directory outside the corpus root), even when the original full run had detected them. Previously the flag was on detect() and collect_files() only. (#736)cluster-only `--graph` flag (#724) — graphify cluster-only now accepts --graph to specify a non-default graph.json location; positional path and flags
--graph flag (#724) — graphify cluster-only now accepts --graph <path> to specify a non-default graph.json location; positional path and flags can appear in any order_is_sensitive() false positives (#718) — word boundaries on the keyword pattern prevent legitimate source files like tokenizer.py, password_verification.py, SecretManager.java from being silently droppedmax_tokens truncation cascade (#730) — headless graphify extract --backend claude/kimi now defaults to 16384 output tokens (was 8192), eliminating the recursive split cascade on dense doc corpora; override with GRAPHIFY_MAX_OUTPUT_TOKENS env var--update now clearly distinguishes "N nodes pruned from M deleted files" from "M deletions detected but graph already clean"source_file (#712) — stub nodes created for imported .svelte files now carry the resolved import path as source_file instead of the importer's pathextract_svelte() now catches static import X from './foo.svelte' via a dedicated regex pass over <script> block content; previously tree-sitter's JS parser silently dropped all static importsgraphify extract (full rebuild path) now saves manifest.json on every successful run; previously only --update saved it, causing stale-manifest driftgraphify antigravity install now writes to .agent/ (no trailing s) matching Antigravity's actual config paths--dedup-llm wiring — flag now correctly threads LLM backend through to deduplicate_entities in both fresh and incremental extract paths; fresh extract path now also runs dedup (previously bypassed it entirely)cluster-only now accepts --graph <path> to specify a non-default graph.json location; positional path and flags can appear in any order (#724)_is_sensitive() no longer drops legitimate source files — word boundaries on the keyword pattern prevent false positives like tokenizer.py, password_verification.py, SecretManager.java (#718)graphify extract --backend claude/kimi raises default max_tokens from 8192 → 16384, eliminating the truncation-then-recursive-split cascade on dense doc corpora; respects GRAPHIFY_MAX_OUTPUT_TOKENS env var (#730)--update prune message now clearly distinguishes "N nodes pruned from M deleted files" from "M deletions detected but graph already clean — no drift" (#539)extract_svelte() stub nodes now carry the resolved import path as source_file instead of the importer's path, preventing metadata corruption after merge (#712)extract_svelte() now catches static import X from './foo.svelte' via a dedicated regex pass over <script> block content — previously tree-sitter's JS parser silently dropped all static imports in .svelte files (#713)graphify extract (full rebuild path) now saves manifest.json on every successful run, not only on --update; prevents stale-manifest drift on subsequent incremental runs (#538)graphify antigravity install now writes to .agent/ (no trailing s) matching Antigravity's actual config paths (#704)--dedup-llm flag now correctly threads LLM backend through to deduplicate_entities in both fresh and incremental extract paths; fresh extract path now also runs dedup (previously called build_from_json directly, bypassing dedup entirely)Feat: graphify extract now runs incrementally — auto-detects prior manifest.json and re-extracts only changed/new files; semantic results cached by co
graphify extract now runs incrementally — auto-detects prior manifest.json and re-extracts only changed/new files; semantic results cached by content hash so unchanged docs cost zero LLM tokens on repeat runs (#698)--dedup-llm flag for graphify extract — optional LLM tiebreaker for ambiguous entity pairs (~$0.01 for 10k-node graphs), off by defaultgraphify hook install rebuild now preserves human-readable community labels from .graphify_labels.json instead of resetting to generic "Community N" names on every commit (#705)graphify install --platform gemini now works correctly (#706)datasketch and rapidfuzz added as base dependencies_read_tsconfig_aliases() now handles the format every TypeScript tool generates — // comments, /* */ block comments, trailing commas. Previously json.
_read_tsconfig_aliases() now handles the format every TypeScript tool generates — // comments, /* */ block comments, trailing commas. Previously json.loads threw JSONDecodeError which was silently swallowed, causing alias resolution to return {} for essentially every real TS project (SvelteKit, NestJS, T3, Astro, Vite, Nx). The extends-following logic from v0.7.1 was correct all along — it was just unreachable. Both now work together.
Two bugs in the extract_svelte() regex fallback that caused 100% of generated edges to be dropped:
$lib/..., $partials/..., @/...) were skipped by a .startswith('.') filter — now resolved via the tsconfig alias map_make_id signature so both endpoints were phantom nodes; fixed to match _extract_generic and _import_js conventionsFor a real 1,870-file SvelteKit app: 4,246 $lib/... + 478 $partials/... imports now resolve correctly, recovering ~150 previously orphaned .svelte files.
_read_tsconfig_aliases() now parses JSONC — handles // line comments, /* */ block comments, and trailing commas that every TypeScript framework starter generates; warns to stderr on parse failure instead of silently returning {} (#700)extract_svelte() regex fallback now captures aliased dynamic imports ($lib/..., $partials/..., @/...) and uses correct _make_id(str(path)) scheme so edges survive into graph.json instead of being dropped as phantom nodes (#701)Docs-only and mixed corpora can now run full LLM extraction without Claude Code:
graphify extract — headless semantic extraction for CI (#698)Docs-only and mixed corpora can now run full LLM extraction without Claude Code:
pip install --upgrade graphifyy
export MOONSHOT_API_KEY=... # or ANTHROPIC_API_KEY
graphify extract ./docs
What it does: detect → AST extract on code files → semantic LLM extract on docs/papers/images → merge → cluster → write graph.json + .graphify_analysis.json.
Flags:
--backend kimi|claude — explicit backend (auto-detected from env if omitted)--out DIR — output root (default: <path>/graphify-out/)--no-cluster — skip community detection, write raw extraction onlySupported backends: Anthropic (ANTHROPIC_API_KEY) and Kimi (MOONSHOT_API_KEY). When using the /graphify skill interactively, any model your IDE session runs is used instead.
graphify extract <path> — headless full-pipeline extraction for CI; runs AST extraction on code files and semantic LLM extraction on docs/papers/images without Claude Code in the loop; supports --backend kimi|claude, --out DIR, --no-cluster; auto-detects backend from MOONSHOT_API_KEY / ANTHROPIC_API_KEY; docs-only corpora (issue #698) work cleanlyParses .f, .F, .f90, .F90, .f95, .F95, .f03, .F03, .f08, .F08 files
.f, .F, .f90, .F90, .f95, .F95, .f03, .F03, .f08, .F08 files*.F, *.F90, etc.) are preprocessed with cpp -w -P before parsing to handle #ifdef/#define directivesNew graphify export <format> subcommands replace the Python snippet workflow:
graphify export html # interactive graph.html
graphify export obsidian # Obsidian vault (--dir PATH)
graphify export wiki # wiki/ markdown articles
graphify export svg # static SVG
graphify export graphml # GraphML for Gephi/yEd
graphify export neo4j # cypher.txt (or --push URI for live import)
Also available: graphify query, graphify path, graphify explain
Reduced from 63KB to 47KB by replacing multi-line Python heredocs with single-line CLI calls. Fixes #696.
to_html() now accepts node_limit — when exceeded, automatically renders a community meta-graph instead of the full graph.
Requires: pip install tree-sitter-fortran for Fortran support (added as base dependency in this release)
use imports, and call edges from .f, .F, .f90, .F90, .f95, .F95, .f03, .F03, .f08, .F08 files; names are lowercased for case-insensitive matching (#694)Three bug fixes for edge cases that were causing silent failures or broken outputs.
Three bug fixes for edge cases that were causing silent failures or broken outputs.
Obsidian tags with special characters (#690)
Community labels containing ., &, (, ) were producing invalid Obsidian tags that broke Dataview queries. Tags are now sanitized to [a-zA-Z0-9_\-/] only before being written to the vault.
SvelteKit / Nuxt / NestJS path aliases (#691)
tsconfig.json path aliases defined in extended configs (via the `extends` field) were silently dropped. The alias resolver now follows the full extends chain, so `@/components` resolves correctly even when the alias lives in a base config two levels up.
Svelte template-layer dynamic imports (#692)
Tree-sitter only parses the script block of .svelte files. Markup-level dynamic imports like `{#await import('./Modal.svelte')}` in the template were invisible to the graph. A regex pass now runs after AST extraction to catch these.
Recursion crash on large codebases (#695)
Files with deeply nested AST trees (generated code, minified output, some auto-generated schemas) could hit Python's default 1000-frame recursion limit and crash the whole extraction run. The limit is now raised to 10,000 at startup and in each worker process. A `_safe_extract` wrapper catches any file that still overflows and skips it with a clear warning — the rest of the run completes normally.
pip install --upgrade graphifyy
or
uvx graphifyy
., &, (, ) now produce valid Obsidian tags; only [a-zA-Z0-9_\-/] characters survive, preventing broken Dataview queries (#690)_load_tsconfig_aliases() now follows tsconfig extends chains - SvelteKit, Nuxt, and NestJS path aliases defined in extended configs are no longer silently dropped (#691).svelte files now get a regex pass over the template layer after JS AST extraction - {#await import('./X.svelte')} markup-level dynamic imports are captured as edges (#692)_safe_extract wrapper that skips pathological files with a clear warning instead of crashing the whole run (#695)One of the most common pain points I kept hearing about: graphify works great on a solo project, but once you have multiple people committing, things
Hey everyone,
One of the most common pain points I kept hearing about: graphify works great on a solo project, but once you have multiple people committing, things get messy. graph.json ends up with conflict markers, two devs rebuild from the same code and get different community IDs, renamed files blow away the cache. This release fixes all of that.
No more merge conflicts in graph.json
Run graphify hook install once in your repo. It now sets up a git merge driver that union-merges two graph.json files automatically. When two teammates commit at the same time and git tries to merge their graphs, it just combines the nodes and edges from both sides. No conflict markers, no manual resolution, nothing to think about.
It also writes the right .gitattributes entry and registers the driver in .git/config for you.
Deterministic community IDs
Leiden community detection is now seeded. Before this, two people rebuilding the graph from identical code could get different community numbers, which meant a 500-node diff for a 2-line code change. Now parallel rebuilds produce the same output.
Renamed files reuse their cache
The file hash used to include the file path, so renaming utils.py to helpers.py would invalidate the cache and re-extract the whole file. Now it's content-only. Rename 50 files in a refactor and the cache still hits.
Graph freshness signal
graph.json now records which git commit it was built from. GRAPH_REPORT.md shows the short hash so your AI assistant (or you) can compare against git rev-parse HEAD and know if the graph is stale before answering architecture questions.
Mixed code/doc commits handled correctly
If a commit touched both Python files and markdown docs, watch mode used to skip the code rebuild entirely. Now it rebuilds the code immediately and queues the docs for LLM re-extraction separately.
pip install --upgrade graphifyy
or
uvx graphifyy
Then run graphify hook install in your repo to set up the merge driver and post-commit hooks.
Multi-dev busy-repo support: four gaps that caused merge conflicts, stale graphs, and silent cache misses in team workflows.
graphify hook install now also configures a git merge driver for graphify-out/graph.json — union-merges two graph.json files so git never produces conflict markers in the knowledge graph; writes .gitattributes and registers graphify merge-driver in .git/configgraphify merge-driver <base> <current> <other> subcommand — takes two graph.json variants and writes their node/edge union back to <current>; always exits 0 so merge never blocksseed=42 when supported) for deterministic community IDs across parallel rebuilds — reduces JSON diff churn in multi-dev reposgraph.json now embeds built_at_commit (git HEAD) at write time; GRAPH_REPORT.md surfaces the commit hash and a freshness check hintfile_hash is now content-only (path removed from hash) — renamed files reuse their cache entry instead of re-extracting; cached source_file fields are updated to the new path on loadneeds_update flag; previously code changes were silently dropped in mixed batchesFix: source_file path separators normalized to forward slashes at graph ingestion — same physical file emitted with backslashes (Windows AST extractor
source_file path separators normalized to forward slashes at graph ingestion — same physical file emitted with backslashes (Windows AST extractor) and forward slashes (semantic subagents) now merges into one node instead of splitting into two disconnected components (#683)CLAUDE.md) from merging unrelated subsystems into one giant community (#683)GRAPH_REPORT.md, explicit trigger list, narrow allowlist for raw source reads (#688)GRAPHIFY_OUT env var overrides the output directory — accepts a relative name or absolute path, wires through cache.py, watch.py, and the CLI; useful for sharing one graph across multiple git worktrees (#686)graphify antigravity install now auto-updates stale rules and workflow files on re-run instead of silently skipping them (#652)docs/how-it-works.mdFix: .graphifyignore negation patterns (!src/) now work correctly — when any ! pattern is present, directory pruning is deferred to per-file checks so
.graphifyignore negation patterns (!src/**) now work correctly — when any ! pattern is present, directory pruning is deferred to per-file checks so negated files inside ignored directories are reached (#676)/graphify now appears in the command dropdown — workflow file now includes YAML frontmatter with name: graphify required for Antigravity discovery (#678)[ -f ... ] && echo (bash-only) with cross-platform python -c using json.dumps — fixes hook failure on Windows CMD and Git Bash (#681)additionalContext rejection on Codex Desktop PreToolUse (#651)graphify install --platform codex now writes absolute path to graphify executable — fixes PATH resolution in VS Code extension on Windows (#651)GRAPH_REPORT.md by default; report header shows (N total, M thin omitted) and Knowledge Gaps collapses thin communities to one summary line (#664)`graphify tree` — self-contained D3 v7 collapsible-tree HTML view of graph.json; expand/collapse controls, depth-based colours, hover inspector
graphify tree — self-contained D3 v7 collapsible-tree HTML view of graph.json; expand/collapse controls, depth-based colours, hover inspector (#557)query_graph tool now accepts context_filter for cross-language edge filtering (#573)import() extraction — JS/TS dynamic imports now appear as imports_from edges (#579)save_semantic_cache crashed with IsADirectoryError when a node's source_file was a directory — changed p.exists() to p.is_file() (#655)sanitize_label(None) raised TypeError, crashing to_html on graphs with rationale nodes that have null source_file (#656)rationale from valid file_type values — model hallucinated concept on every doc/paper run; all skill variants updated (#657)cost.json always reported 0 tokens — added explicit chunk-merge step that globs and sums real token counts across all skill variants (#658)pip install --upgrade graphifyy
uvx graphifyy install
graphify tree — self-contained D3 v7 collapsible-tree HTML view of graph.json; expand/collapse controls, depth-based colours, hover inspector; XSS-safe (#557)query_graph tool (#573)import() extraction for JS/TS (#579)save_semantic_cache crashed with IsADirectoryError when a node's source_file was a directory path — p.exists() → p.is_file() (#655)sanitize_label(None) raised TypeError crashing to_html on graphs with null source_file rationale nodes — return "" early (#656)rationale from valid file_type values — model hallucinated concept on every doc/paper run; explicit merge step added to all skill variants (#657)cost.json always reported 0 tokens — chunk JSONs have placeholder zeros; orchestrator now globs and sums real token counts before merging (#658)graphify pi install / graphify pi uninstall — installs skill to ~/.pi/agent/skills/graphify/SKILL.md
graphify pi install / graphify pi uninstall — installs skill to ~/.pi/agent/skills/graphify/SKILL.mdgraphify install --platform pi also supportedskill-windows.md) — complete rewrite from PowerShell to bash. Claude Code on Windows uses git-bash, so PowerShell syntax (\$null, \$LASTEXITCODE, Select-Object, Remove-Item) caused exit code 49 failures. Now mirrors skill.md exactly, with python added as fallback after python3 for Windows Conda environments (#39)to_wiki() now clears old .md files before regenerating, preventing orphan article accumulation across runs (#558)_safe_filename() now strips all Windows-reserved characters (< > : " / \ | ? *) and caps length at 200 chars (#594)extract.py (#576)calls and rationale_for edges were being flipped in graph.json and graph.html due to NetworkX undirected storage canonicalizing endpoint order. True direction now restored from _src/_tgt metadata stashed by build.py (#576)skill-trae.md SyntaxError — stray colon in --cluster-only block caused install failures (#603)detect.py marked usedforsecurity=False (used for file change detection only, not cryptography)transcribe.py marked usedforsecurity=False (used for filename generation only)pyproject.toml: tightened wheel packaging to only include the graphify package — wheel drops from 1.7MB to 285KB.graphifyinclude hidden path allowlist — opt specific hidden files/dirs into traversal (e.g. .hermes/plans/) (#583)--no-viz flag now works in cluster-only path — removes stale graph.html when passed (#565)GRAPHIFY_VIZ_NODE_LIMIT env var — override the 5000-node HTML threshold (set to 0 to disable viz entirely, useful for CI) (#565)--update prune output now splits into drift-detected vs no-drift cases for clarity (#544)--update merge step now calls save_manifest to prevent deleted files reappearing as ghost nodes on subsequent runs (#545)pip install graphifyy==0.6.6
# or
uv tool install graphifyy
skill-windows.md rewritten from PowerShell to bash — Claude Code on Windows uses git-bash so PowerShell syntax ($null, $LASTEXITCODE, Select-Object, & (Get-Content ...), Remove-Item) caused exit code 49 failures; now mirrors skill.md structure with python added as fallback after python3 for Windows Conda (#39)to_wiki() now clears stale articles before regenerating, preventing orphan .md accumulation (#558)_safe_filename() in wiki.py now strips Windows-reserved characters (< > : " / \ | ? *) and caps length at 200 chars (#594)calls, rationale_for) preserved correctly at JSON export (#576).graphifyinclude hidden path allowlist — opt specific hidden dirs into traversal (e.g. .hermes/plans/**/*.md) (#583)--no-viz flag wired in cluster-only; GRAPHIFY_VIZ_NODE_LIMIT env var overrides 5000-node HTML threshold (#565)skill-trae.md --cluster-only block (#603)--update prune output clarified — splits no-drift vs drift cases (#544)--update merge step now calls save_manifest to prevent deleted files reappearing (#545)graphify tree — self-contained D3 v7 collapsible-tree HTML view of graph.json; expand/collapse controls, depth-based colours, hover inspector; XSS-safe via html.escape() and _js_safe() (#557)Fix: Codex PreToolUse hook on Windows (#651, #522) — replaced python3 -c "..." (breaks on Conda, broken by PowerShell JSON parsing) with graphify hook
python3 -c "..." (breaks on Conda, broken by PowerShell JSON parsing) with graphify hook-check, a new shell-agnostic subcommand. Re-run graphify codex install to regenerate the hook.simple_identifier and identifier node types across grammar versionsgraphify update --force (#639) — bypass node-count safety check after refactors that legitimately shrink the graph; also GRAPHIFY_FORCE=1 env varpip install --upgrade graphifyy
Windows Codex users: also re-run graphify codex install after upgrading to regenerate the hook.
simple_identifier and identifier node types — PyPI's tree_sitter_kotlin grammar uses identifier while older forks use simple_identifier, causing zero calls edges to be emitted (#659)graphify update --force and GRAPHIFY_FORCE=1 env var — bypass the node-count safety check after refactors that legitimately shrink the graph (#639)python3 -c "..." inline command (fails on Conda where only python exists, and breaks PowerShell JSON parsing) with graphify hook-check, a new shell-agnostic subcommand. Re-run graphify codex install to regenerate the hook (#651, #522)Codex PreToolUse hook fails on Windows (#651) — the hook used [ -f graphify-out/graph.json ] which is bash-only and crashes on cmd.exe. Replaced with
[ -f graphify-out/graph.json ] which is bash-only and crashes on cmd.exe. Replaced with a cross-platform Python one-liner (pathlib.Path.exists()) that works on Windows, Linux, and macOS.pip install --upgrade graphifyy
# or
uvx graphifyy install
After upgrading, re-run graphify codex install to regenerate the hook with the fixed command.
[ -f ] is bash-only and crashes on cmd.exe; replaced with a cross-platform Python one-liner (pathlib.Path.exists()) (#651)Incremental rebuild drops semantic nodes (#653) — graphify update and post-commit hooks now preserve INFERRED/AMBIGUOUS nodes by ID membership rather
graphify update and post-commit hooks now preserve INFERRED/AMBIGUOUS nodes by ID membership rather than file_type, so LLM-extracted call/data-flow edges survive code-only rebuildsnohup & disown, git returns in ~100ms, rebuild log at ~/.cache/graphify-rebuild.logcalls now skips any callee resolving to 2+ candidates (ambiguous short names like log, execute, find no longer accumulate hundreds of spurious edges)to_html in cluster command now guarded with try/except ValueError matching the watch/hook pathpip install --upgrade graphifyy
# or
uvx graphifyy install
graphify update, post-commit hook) dropped INFERRED/AMBIGUOUS semantic nodes extracted from code files — node preservation now filters by ID membership in the new AST output instead of file_type, so LLM-extracted call/data-flow edges survive code-only rebuilds (#653)git commit for the full rebuild duration (hours on large repos) — rebuilds now detach via nohup & disown, git returns in ~100ms, log written to ~/.cache/graphify-rebuild.log (#650)calls resolution used a last-write-wins name map, causing common short names (log, execute, find) to accumulate hundreds of spurious edges and dominate god_nodes ranking — resolution now skips any callee name that matches 2+ candidates (ambiguous, no import evidence to pick the right target) (#543)cluster-only command crashed on graphs with >5000 nodes due to unguarded to_html call — now wrapped in try/except ValueError matching the watch/hook path (#541)Query exact-match: graphify query "MyFunction" now returns the actual function first, not hub modules
graphify query "MyFunction" now returns the actual function first, not hub moduleshtml2text (GPL-3.0) with markdownify (MIT) — no more copyleft in a MIT projectgraphify update never persisted manifest — every run re-extracted everything (#621).graphifyignore (vendor/ # legacy) now stripped correctly (#605).tmp cache file — unique tempfile per writer (#589)- could inject git flags in _clone_repo (#589)html2text (GPL-3.0) with markdownify (MIT) (#586)cluster-only CLI flipped directed graphs to undirected (#590).graphifyignore negation patterns (!pattern) now work with last-match-wins (#628).r) support via LLM semantic extraction (#617)pip install --upgrade graphifyy
content empty — thinking now disabled on Moonshot calls so graphs actually populate (#623)graphify update / graphify watch never persisted the manifest, so every subsequent --update re-extracted all files — manifest now saved after each rebuild (#621).graphifyignore (e.g. vendor/ # legacy) now stripped correctly — whitespace + # suffix is treated as a comment, path#hash.py preserved (#605)graphify query "FunctionName" now returns the exact matching node first instead of high-degree hub modules hijacking the output — 100-point exact-match bonus + seeds render before BFS expansion (#638).tmp cache file — each writer now gets a unique tempfile via mkstemp, eliminating cache corruption under parallel extraction (#589)_clone_repo branch names starting with - could be interpreted as git flags — validation added, -- separator inserted before positional args (#589)html2text (GPL-3.0) with markdownify (MIT) — removes the only copyleft dependency from a MIT project (#586)--update re-extracted files whose mtime was bumped by sync tools (Obsidian, Nextcloud) without content changes — manifest now stores content hash alongside mtime; mtime bump triggers an MD5 check before re-extraction (#593).r files classified as code and processed via LLM semantic extraction (#617)#!/bin/bash, #!/usr/bin/env python3, etc.) and included as code (#619)calls edges (e.g. Python→TypeScript name collision) no longer appear as top surprising connections in GRAPH_REPORT.md (#630)cluster-only CLI silently flipped directed graphs to undirected — directed flag now read from graph.json and preserved through re-clustering (#590)\\?\C:\...) now normalize to consistent cache keys (#629).graphifyignore negation patterns (!src/lib/secrets.ts) now work — full last-match-wins evaluation with ! un-ignore support (#628)Correct gitignore semantics: outer rules load first so inner (closer) rules always win via last-match-wins — matching standard gitignore behavior exac
.graphifyignore discovery stops at the scan folder — no leakage across sibling projects in a shared workspace (#643)/ patterns in a parent .graphifyignore now apply only relative to their own directory, not the scan root (#643)vendor\ (escaped) is preserved per gitignore spec (#643)``` pip install --upgrade graphifyy ```
.graphifyignore discovery now uses correct gitignore semantics — outer rules are loaded first so inner (closer) rules always win via last-match-wins, matching standard gitignore behavior (#643).graphifyignore discovery is now hermetic to the scan folder — no leakage across sibling projects in a shared workspace (#643)/) in a parent .graphifyignore now correctly apply only relative to their own directory, not the scan root (#643)vendor\ (escaped) is preserved (#643).sql files are now processed deterministically via tree-sitter — no LLM needed, no tokens spent.
.sql files are now processed deterministically via tree-sitter — no LLM needed, no tokens spent.
graphify extracts:
CREATE TABLE → table nodesCREATE VIEW → view nodesCREATE FUNCTION / CREATE PROCEDURE → function nodesREFERENCES constraints → foreign key edges (EXTRACTED confidence)FROM / JOIN clauses → reads_from edges (EXTRACTED confidence)This means your entire backend — application code + database schema — maps into one graph. Tables, views, stored functions, and their relationships are first-class nodes just like Python classes or Go structs.
Supports standard SQL, PostgreSQL, Snowflake, and any dialect the tree-sitter-sql grammar covers.
Setup: ``` pip install 'graphifyy[sql]' /graphify . ```
.yaml and .yml files are now picked up for semantic extraction. Kubernetes manifests, Kustomize overlays, Helm values, and any YAML config are indexed automatically — no extra setup needed.
NameError: _os is not defined crash after graphify update.graphify_labels.json now persists across re-clustersSyntaxWarning in shell glob pattern``` pip install --upgrade graphifyy pip install 'graphifyy[sql]' # for SQL support ```
.sql files now processed deterministically via tree-sitter. Extracts tables, views, functions/procedures, foreign key references, and FROM/JOIN reads_from edges. No LLM needed. Requires pip install 'graphifyy[sql]' (#349)xlsx_extract_structure() utility — extracts sheet names, named tables, and column headers from .xlsx files as structural nodesFeat: YAML/YML files now indexed for semantic extraction — Kubernetes, Kustomize, Helm, and any YAML corpus now picked up automatically
Fix: NameError: name '_os' is not defined crash after graphify update — this was fixed in v5 branch but not released to PyPI (#618, #612)
NameError: name '_os' is not defined crash after graphify update — this was fixed in v5 branch but not released to PyPI (#618, #612)Kimi K2.6 backend — pip install 'graphifyy[kimi]' then set MOONSHOT_API_KEY to route semantic extraction through Kimi K2.6 instead of Claude subagents
pip install 'graphifyy[kimi]' then set MOONSHOT_API_KEY to route semantic extraction through Kimi K2.6 instead of Claude subagents. 3-6x richer relation extraction at ~3x lower cost. Claude remains the default; Kimi is opt-in. A tip is printed when the key is not set so users discover it naturally.this.logger.log() → log) are no longer cross-file resolved. Go package-qualified calls (pkg.Func()) are correctly preserved. Affects JS/TS, Go, Rust, Swift, Kotlin, Scala, PHP, C++, C#, Zig, Elixir.file_type: concept no longer produce validation warnings.graphify-out/.graphify_root so graphify update with no args finds the right directory automatically.Fixes #598, #601.
pip install 'graphifyy[kimi]' + MOONSHOT_API_KEY routes semantic extraction through Kimi K2.6. 3-6x richer relation extraction at ~3x lower cost. Claude remains default; Kimi is opt-in.this.logger.log() → log) no longer cross-file resolved. Go package-qualified calls (pkg.Func()) correctly preserved. Affects JS/TS, Go, Rust, Swift, Kotlin, Scala, PHP, C++, C#, Zig, Elixir.concept file_type no longer triggers validation warnings (#601)graphify update remembers scan root via graphify-out/.graphify_root — no path argument needed on subsequent runs.graphify_labels.json is now preserved so wiki/obsidian/HTML retain human-readable names after re-cluster (#608)NameError: name '_os' is not defined in graphify update Kimi tip (#612)SyntaxWarning in __main__.py for shell glob pattern with backslash escapesrequires-python = ">=3.10" now supports Python 3.14+ (#607)SSRF DNS rebinding fix — safe_fetch now patches socket.getaddrinfo for the entire duration of each HTTP request so a DNS rebinding attack cannot swap
safe_fetch now patches socket.getaddrinfo for the entire duration of each HTTP request so a DNS rebinding attack cannot swap a public IP (returned during validation) for a private one during the actual connection. DNS lookup failures now also raise an error instead of silently skipping the IP check.download_audio now runs validate_url before handing the URL to yt-dlp, blocking private IPs and disallowed schemes on the video/audio ingest path.Fixes #591, #592.
safe_fetch now patches socket.getaddrinfo for the full request duration (#591)download_audio now calls validate_url before handing URL to yt-dlp (#592)Fixes `graphify update` broken on mixed code+docs corpora
Fixes graphify update broken on mixed code+docs corpora (#582)
AST and semantic cache entries now live in separate cache/ast/ and cache/semantic/ subdirectories. Previously both wrote to the same flat cache/ directory — semantic results silently overwrote AST entries for code files, causing the shrink guard to fire on every subsequent graphify update run.
Existing flat cache entries are read as a migration fallback so no cache is lost on upgrade.
Upgrade: pip install --upgrade graphifyy
cache/ast/ and cache/semantic/ subdirectories; flat entries read as migration fallbackHotfix for Claude Code v2.1.117+
Hotfix for Claude Code v2.1.117+
The PreToolUse hook installed by graphify claude install was silently broken on Claude Code v2.1.117+ because dedicated Grep and Glob tools were removed — searches now go through Bash. The hook never fired, meaning the graph context reminder was never injected.
Fix: Hook matcher changed from Glob|Grep to Bash. The hook now reads the tool input JSON from stdin and pattern-matches on the command string, firing only on search-like calls (grep, rg, find, fd etc.) — not on every shell command.
Upgrade: pip install --upgrade graphifyy && graphify claude install
Closes #578
Bash instead of Glob|Grep for Claude Code v2.1.117+6 bug fixes shipped in this release:
6 bug fixes shipped in this release:
utils.py files) now get unique node IDs via parent-directory-qualified stemssource_file paths — extract() relativizes all paths before returning; graph.json is now portable across machines and git worktreesto_json() returns a boolean; graphify update only writes GRAPH_REPORT.md and graph.html if the JSON write succeeded (shrink guard fired = no stale report)@/ and other compilerOptions.paths aliases in tsconfig.json now resolve to real file nodes instead of being dropped as external packagescalls edge direction explicitly enforced (caller → callee)Also includes tooling fixes from the v0.5.0 patch: ~ expansion in core.hooksPath, correct .gitignore inline comment placement, # nosec annotations on file write sinks.
source_file paths relativized before return so graph.json is portableto_json() returns bool; report only written on successful JSON write@/ path aliases resolved via tsconfig.jsonClone any GitHub repo directly `bash graphify clone https://github.com/karpathy/nanoGPT Clones to ~/.graphify/repos/ / , reuses existing clones on rep
Clone any GitHub repo directly
graphify clone https://github.com/karpathy/nanoGPT
Clones to ~/.graphify/repos/<owner>/<repo>, reuses existing clones on repeat runs. Supports --branch and --out.
Cross-repo knowledge graphs
graphify merge-graphs repo1/graphify-out/graph.json repo2/graphify-out/graph.json
Every node carries a repo tag so you can filter by origin.
Data-loss protection on --update
graphify now refuses to overwrite graph.json with a smaller graph. New build_merge() library function for safe incremental updates that only ever grows the graph.
Duplicate node deduplication
Chunk-suffix contamination (_c2, _c4) blocked at the prompt level. Post-merge deduplication pass catches any stragglers.
Bug fixes
CLAUDE_CONFIG_DIR env var respected on install (#527)graphify-out/ excluded from source scanning (#524)Full announcement: https://github.com/safishamsi/graphify/discussions/528
graphify clone <github-url> — clone and graph any public repographify merge-graphs — combine multiple graph.json outputs into one cross-repo graphCLAUDE_CONFIG_DIR support in graphify installto_json() refuses to overwrite with a smaller graphbuild_merge() for safe incremental updatesdeduplicate_by_label()graphify-out/ excluded from source scanningThe skill was placing graphify-out/ in the current working directory instead of inside the target path. Fixed by adding cd INPUT_PATH as the very firs
/graphify <path> now respects the specified directoryThe skill was placing graphify-out/ in the current working directory instead of inside the target path. Fixed by adding cd INPUT_PATH as the very first command in Step 1, and changing detect(Path('INPUT_PATH')) → detect(Path('.')) so all relative paths resolve correctly from the target directory. Applied to both skill.md and skill-opencode.md.
_hooks_dir() was using git config core.hooksPath which doesn't resolve the shared hooks directory for worktree checkouts. Now uses git rev-parse --path-format=absolute --git-path hooks which git resolves correctly for both normal repos and worktrees. Includes regression tests for install/uninstall from a worktree.
pyproject.toml already points to the correct repository. GitHub's dependency graph cache will update itself.
## Features & Fixes - #488 Legacy edge/node schema canonicalization — source→source_file on nodes, from/to→source/target on edges now handled before v
source→source_file on nodes, from/to→source/target on edges now handled before validationextends, implements, and interface-extends edges now extracted from AST; two-pass cross-file import resolver emits real imports edges instead of silently dropping themgraph.html instead of no output at all; updated in both skill.md and skill-opencode.mdgraphify check-update <path> — cron-safe CLI subcommand that checks for pending semantic updates and notifies without running LLM extraction441 tests passing.
Added test_to_canvas_file_paths_relative_to_vault — would have caught the hardcoded path bug
head -1 for .exe binaries, goes straight to python3 fallbackto_canvas() hardcoded graphify/obsidian/ path — file refs are now vault-root-relative ({fname}.md), canvas opens correctly regardless of where the vault livestest_to_canvas_file_paths_relative_to_vault — would have caught the hardcoded path bugtest_hook_skips_head_on_exe — verifies the Windows fix is present in the hook script## Fix - #504 Garbled Chinese (CJK) output on Windows — added export PYTHONUTF8=1 to Step 1 of skill-opencode.md so all Python subprocesses run in UTF
export PYTHONUTF8=1 to Step 1 of skill-opencode.md so all Python subprocesses run in UTF-8 mode on Windows, preventing garbled non-ASCII characters in terminal output and intermediate filesopencode.json now written to .opencode/opencode.json instead of project root — consistent with OpenCode's config directory
opencode.json now written to .opencode/opencode.json instead of project root — consistent with OpenCode's config directory.graphify_detect.json, .graphify_extract.json, etc.) in skill-opencode.md now correctly use graphify-out/ prefix instead of leaking to project root--wiki step (Step 6b) to skill-opencode.md, matching all other platform skillsanalyze.py (#499): Add seed=42 to betweenness_centrality() — GRAPH_REPORT.md is now byte-identical across runs on graphs >1000 nodes, eliminating infi
seed=42 to betweenness_centrality() — GRAPH_REPORT.md is now byte-identical across runs on graphs >1000 nodes, eliminating infinite commit churn from the post-commit hookgraph.json edge endpoints are stable across machineswiki.py (#496): Add encoding="utf-8" to all write_text() calls — fixes crash on Windows cp1252 when articles contain Unicode characters like →
encoding="utf-8" to all write_text() calls — fixes crash on Windows cp1252 when articles contain Unicode characters like →_unique_slug() — prevents silent article overwrites when two communities get the same labelgit rebase, git merge, and git cherry-pick — prevents unstaged changes blocking --continueroot path at detect() entry — fixes .graphifyignore patterns from parent directories not matching when graphify is run with a relative path like ./rawmanifest.json and cost.json gitignore recommendations; add .graphifyignore example for platform install filesreport.py: Complete empty-community fix — Community Hubs navigation now skips empty-community entries, Summary line reflects non-empty community count
cache: skip directory source_file in save_cached — prevents IsADirectoryError crash on NestJS/TS monorepos
source_file in save_cached — prevents IsADirectoryError crash on NestJS/TS monorepos (#448)Nodes (0): sections from reports (#443)@ in python path allowlist — fixes silent git hook failures on macOS with Homebrew Python (#473)source_file paths project-relative after code rebuild — fixes absolute paths leaking into graph.json (#434).graphify_chunk_*.json temp files at end of run (#464)translations/ subdirectory`bash pip install --upgrade graphifyy `
graphify install when multiple platforms were previously installed. graphify install now refreshes .graphify_version in all known skill dirs so the warning clears across the board..html files silently skipped during detection. Added .html to DOC_EXTENSIONS — HTML pages, docs, and web content now indexed._rebuild_code (watch/update/git hook) failed entirely on graphs > 5000 nodes because to_html raised ValueError. Wrapped in its own try/except so graph.json and GRAPH_REPORT.md always land; stale graph.html from a previous smaller run is removed."context") produced imports_from edges pointing at local files of the same basename, creating false cycle-dependency pairs. Go import node IDs now prefixed go_pkg_ using the full import path.pip install --upgrade graphifyy
graphify install when multiple platforms were previously installed — graphify install now refreshes .graphify_version in all other known skill directories so the warning clears across the board (#178).html files silently skipped during detection — added .html to DOC_EXTENSIONS; HTML pages, docs, and web project content now indexed correctly (#260)_rebuild_code (watch/update/hook) fails entirely on graphs > 5000 nodes because to_html raises ValueError — wrapped in its own try/except so graph.json and GRAPH_REPORT.md always land; stale graph.html from a previous smaller run is removed (#432)"context") produced imports_from edges pointing at local files of the same basename — Go import node IDs now prefixed go_pkg_ using the full import path, eliminating false cycle-dependency pairs (#431)`bash pip install --upgrade graphifyy `
src/graphify-out/cache/ instead of project root when all code files share a common prefix (e.g. src/). extract() now called with explicit cache_root in _rebuild_code and the Codex skill AST step..mdx files silently skipped during detection. Added .mdx to DOC_EXTENSIONS — Next.js, Docusaurus, and Astro corpora now indexed correctly.pip install --upgrade graphifyy
src/graphify-out/cache/ instead of project root when all code files share a common prefix like src/ — extract() now called with explicit cache_root=watch_path in _rebuild_code and cache_root=Path('.') in the Codex skill AST step (#429).mdx files silently skipped during detection — added .mdx to DOC_EXTENSIONS in detect.py; MDX-based corpora (Next.js, Docusaurus, Astro) now indexed correctly (#428)`bash pip install --upgrade graphifyy `
graphify cluster-only crashed with KeyError: 'total_files' in report.py. Cluster-only skips detection so the stats dict was empty — now passes a warning key so the report skips the file-stats section gracefully./graphify --update dropped all existing graph nodes. The merge block built a correct in-memory merged graph but never wrote it back to .graphify_extract.json, so Step 4 rebuilt from the new-extraction-only file. Merged result is now serialized back before Step 4 runs.pip install --upgrade graphifyy
graphify cluster-only crashed with KeyError: 'total_files' in report.py — cluster-only skips detection so the stats dict was empty; now passes a warning key so the report skips the file-stats section (#422)/graphify --update dropped all existing graph nodes — the merge block built a correct in-memory G_existing but never wrote it back to .graphify_extract.json, so Step 4 rebuilt from the new-extraction-only file; merged result is now serialized back before Step 4 runs (#423)Your coding agent can read these notes before it upgrades. Set up the MCP server →