A C0 control character in a node label or id (from a mangled extraction or an odd
filename) crashed the GraphML export (XML 1.0 forbids C0 except tab/LF/CR) and the
Obsidian export (EINVAL on Windows paths), aborting the whole artifact. Fold a
control-char scrub into the existing per-format coercion hooks so only the illegal
characters are stripped; tab/LF/CR and non-ASCII letters are preserved, and to_json
(and its byte-identity round-trip) is untouched.
The docstring documents the `\w` regex class; as a non-raw string that is an
invalid escape sequence (DeprecationWarning now, SyntaxError in a future Python).
Mark it raw.
_obsidian_tag stripped every non-ASCII character, so a Korean or Japanese community label collapsed to underscores and every note in that community carried the same tag. Python's \w is Unicode-aware, so switching the filter keeps letters from any script while still dropping spaces and punctuation.
Also route the .obsidian/graph.json colour-group query through the same sanitizer: it was built from the raw label, so on a non-ASCII label it queried a tag that no note carries.
node_link_data always appends the node key (id) last, so id sat mid-dict on a
cold build but last after a read-rebuild — churning graph.json key order with
no content change and defeating byte-stable caching. to_json now canonicalizes
each node/link dict (id / source,target,relation first, remaining keys sorted)
before serialization, so a load -> rebuild -> write round-trip is byte-identical.
Pure key permutation; values are untouched.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
AST-emitted INFERRED edges landed at the rubric-forbidden 0.5 (or a hardcoded
0.8). Each AST INFERRED emission site now sets a discrete rubric score keyed to
the relation (uses -> 0.95 direct structural evidence, indirect_call and
unresolved cross-file calls -> 0.85), and the INFERRED write-time default moves
0.5 -> 0.55 so any score-less INFERRED edge is on the rubric set. EXTRACTED /
AMBIGUOUS tiers are unchanged; uses stays INFERRED (not promoted to EXTRACTED)
to keep audit percentages honest.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The semantic extractor emits hyperedges with no id and build.py persists them
verbatim, so a prior graph.json can contain id-less hyperedges. Seeding the
dedup set with a hard h["id"] raised KeyError: 'id' on every incremental
re-extract, silently failing the whole graph load. Guard the comprehension with
h.get("id"), symmetric with the incoming-set guard already below it; id-less
entries are retained in the graph, id-bearing dedup is unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Widens _SLUG_SUFFIX_RESERVE/_DEDUP_SUFFIX_RESERVE from 4 to 5 so a four-digit
collision suffix (_1000..) can't push a truncated stem past MAX_PATH; the
suffix is technically unbounded but 5 chars covers ~10k identical stems. Adds
an end-to-end test that CJK labels at a tight budget stay within the window,
keep their non-ASCII characters, and produce links that resolve on disk.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds paths.stem_filename_budget(output_dir, *, reserve, limit=200) and threads
it through the Obsidian and wiki exporters so a filename stem is budgeted
against the whole Windows MAX_PATH window (drive + dirs + name + NUL), not
just the per-component 200-char NAME_MAX cap. On POSIX the helper returns the
limit unchanged, so existing vaults stay byte-identical; on Windows a long
output directory no longer pushes the total path over MAX_PATH and aborts the
export mid-write.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
#2486 (thanks @adminwat): normalize dict-shaped hyperedge members to ids
(or drop with a warning) so a malformed hyperedge can't abort a completed
merge with a TypeError.
#2484 (thanks @sortakool; approach from @oleksii-tumanov's #1691):
merge-graphs relabels hyperedge member ids and ids with the repo prefix,
unions both inputs' hyperedges instead of clobbering, and writes both
persistence slots.
#2485 (thanks @sortakool): build_from_json reads hyperedges from the
top-level and nested slots; a full validation wipeout is reported loudly.
#2490 (thanks @PapiScholz): the skill Step-5 flow passes curated
community_labels to to_json, so graph.json ships community_name.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Follow-ups on the cherry-picked #2242/#2232: an all-dots label ('...') no
longer produces an empty 'dot-' Obsidian stem (falls back to 'unnamed'),
and the .env.example carve-out gets the regression test it shipped without
(templates graphable, real .env still sensitive, secrets/.env.example still dropped).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
safe_name left stems like .env intact, so the vault wrote .env.md which
Obsidian treats as a hidden file — invisible in the explorer and as
unresolved wikilinks. Prefix with dot- (shared _obsidian_safe_stem for
vault + canvas). True label stays in the note body.
Fixes#2205
graph.json (the clustered `to_json` write and the `--no-cluster`/merge raw
dumps) and manifest.json were written with a direct `open()`/`write_text`, so a
crash, kill, or disk-full mid-write left a truncated, unparseable file that the
next load or `detect_incremental` then failed on.
Add `write_text_atomic`/`write_json_atomic` in graphify.paths (temp file in the
same directory + `os.replace`; JSON is streamed into the temp, not materialized
as one string) and route the graph.json writers (export.to_json,
cli._prune_graph_json_sources, the merge driver) plus detect.save_manifest
through them. The helper preserves the destination's mode (an atomic replace
never tightens 0644 to mkstemp's 0600), writes through a symlinked destination
(shared-output setups), and falls back to copy-then-delete on a Windows
os.replace lock — matching graphify.cache's existing atomic writer. On failure
the previous file is left intact and the temp removed. Not a power-loss
durability guarantee (no fsync, consistent with the rest of the codebase).
Two gaps the review found in the incomplete-build shrink guard:
- existing_graph_node_count() returned None ("proceed") on a present-but-
unparseable graph.json, so the --no-cluster path could clobber a complete
graph whose file was corrupt/mid-write. It now returns a MALFORMED_GRAPH
sentinel and the caller fails closed, matching to_json's #479 handling.
- A walk that couldn't fully enumerate the corpus (permission-denied subtree,
I/O error) is now treated as an incomplete extraction: detect()/
detect_incremental() already record walk_errors; the extract path consumes
them so a walk-truncated graph can't force-overwrite a complete one.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The clustered write is guarded by to_json's #479 shrink check, but the
`--no-cluster` raw-dump path writes graph.json directly and had no guard, so an
incomplete `--no-cluster` build could still overwrite a larger complete graph
with a partial one — the residual gap noted in the original change.
Add `existing_graph_node_count` in graphify.export (mirrors to_json's guard and
respects the graph-size cap) and, on the raw path, refuse the write with exit 1
before the manifest when the build was incomplete and the new graph has fewer
nodes than the existing one — unless --allow-partial is passed. Both write paths
now enforce the same guarantee.
Re-exporting into an existing vault left notes for nodes that dropped
out of the graph, and rewriting the manifest to only this run's files
disowned those orphans - so a returning node's stale note became
permanently unwritable.
Before rewriting the manifest, delete stale = owned - written - skipped.
This only ever touches files graphify itself wrote (foreign files go to
_skipped, never the manifest), and each path is containment-checked
against the vault dir to defuse a corrupt/hostile manifest with ../
entries. The manifest rewrite then correctly drops the pruned files.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
#1831 — `graphify export graphml` crashed on any dict/list-valued
attribute (per-node metadata dict, graph-level hyperedges list) because
nx.write_graphml only accepts scalars; a real ~2,300-node graph failed
every export and left a 0-byte .graphml behind. to_graphml now coerces
None->"" and JSON-serializes non-scalars across graph/node/edge scopes
(int/float/bool/str pass through), and writes atomically via a temp file
so a failed export can't leave a partial file. Closes#1830.
#1807 followup — adopt @varuntej07's explicit in-guard sys.stdout.flush()
from #1811: piped stdout is block-buffered, so a small fully-buffered
output would only flush at interpreter shutdown (outside the guard),
where a closed-pipe reader escapes as a noisy shutdown error and nonzero
exit. Flushing inside the try closes that gap. Closes#1811.
Reported by @hofmockel (#1831) and @varuntej07 (#1807/#1811).
Co-Authored-By: hofmockel <hofmockel@users.noreply.github.com>
Co-Authored-By: varuntej07 <varuntej07@users.noreply.github.com>
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two correctness fixes found while analysing the reported 'graphify update
occasionally writes a partial graph.json' bug.
Enumeration (P0): detect()'s os.walk had no onerror handler, so any os.scandir
failure -- a transient PermissionError, or a directory created/deleted mid-walk
by concurrent writes (e.g. benchmarking racing the scan) -- was silently
swallowed and that entire subtree dropped out of the file list with no log, no
error. Downstream that becomes a silently partial graph.json. The walk now
records each skipped directory (surfaced as walk_errors in detect()'s result)
and warns to stderr, while still enumerating the rest of the tree. This stays
visible even when a --force/GRAPHIFY_FORCE rebuild bypasses the shrink guards.
Relatedly, to_json's #479 anti-shrink guard was fail-OPEN: a non-empty but
unreadable existing graph.json (corrupt or mid-write) proceeded with the
overwrite. It now fails SAFE -- refuse and point at force=True -- while an
empty/whitespace existing file (no nodes to lose) still proceeds. The size-cap
check keeps running before any read, so an oversized existing file is not
loaded into memory.
Pascal edges (P1): a class method declared in the interface section and defined
in the implementation section each emitted a "method" edge to the same node id,
and the edge helpers (unlike the node helpers) did not dedup, so ~half of a
Pascal/Delphi graph's method edges were doubled -- inflating degree/centrality
and tripping the #1739 cross-file resolver's single-owner god-node guard. Both
extractors now dedup edges on (source, target, relation).
Adds regression tests for all three behaviours.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
export.py mixed nine independent export targets in one 1,671-LOC module. Move
the two largest self-contained ones into a new graphify/exporters/ package,
verbatim, re-exported from export.py so every importer (from graphify.export
import to_html / push_to_neo4j) is unchanged:
- exporters/html.py: to_html + its private helpers (_html_script, _html_styles,
_hyperedge_script, _viz_node_limit) and MAX_NODES_FOR_VIZ.
- exporters/graphdb.py: push_to_neo4j, push_to_falkordb.
- exporters/base.py: COMMUNITY_COLORS (shared by the HTML/SVG/Obsidian
exporters), homed here so per-format modules and export.py both import it
without a cycle.
Verified: AST closure-privacy analysis, facade object identity, ruff clean.
export.py drops 1,671 -> 962 LOC. Full suite unchanged: 3036 passed, 29 skipped.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The #1236 fix guarded to_obsidian's member loop but not to_canvas, so
`graphify export obsidian` (which also writes graph.canvas) still crashed with
KeyError on a community member id absent from G — after the notes exported,
leaving a partial mirror. Reported on 0.9.5 by @swells808.
Apply the same `m in G and m in node_filenames` filter in both to_canvas loops:
the box-sizing loop (so the group box matches the cards actually laid out) and
the card-layout loop (so the sort/label deref and the node_filenames fallback
never touch a dangling id). Regression test added alongside the to_obsidian one.
Full suite 2872.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A hyperedge's member list is canonically keyed `nodes`, but producers
(LLM/subagent drift, externally-supplied graph.json) sometimes emit
`members` or `node_ids` — graphify only read `nodes`, so those hyperedges
silently lost their members, and semantic_cleanup's prune dropped them
entirely. Normalize the member key to `nodes` at one ingest chokepoint in
build_from_json (and in semantic_cleanup, which runs pre-build), deduping
and warning, so every downstream consumer sees the canonical key. Mirrors
the existing from/to edge-endpoint aliasing.
Reported by @askalot-io.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Projects the verdicts `graphify reflect` already distills (preferred /
tentative / contested, exponential time-decayed) into a derived
experiential layer the read surfaces consume, so accumulated agent
experience actually shows up where you look — without polluting the
structural graph.
Design (grounded in agent-memory + provenance literature; a redesign of
the #1542 approach):
- SIDECAR, not graph.json stamping. `reflect` writes `.graphify_learning.json`
next to graph.json (an additional output, so the git hooks produce it
automatically). graph.json stays purely structural; nothing leaks into
GraphML; no graph.json churn. Mirrors the named-graph / event-sourcing
separation of durable truth from a derived layer.
- Reuses the existing reflect aggregate (its `_decay` is the
recency-weighted exponential model; `_finalize_sources` the
classification) — no new scoring.
- PROVENANCE: each verdict carries the source questions/dates that produced
it (cap 5, most-recent first).
- STALENESS: each verdict stores the node's file fingerprint; on read, a
changed source file flags the verdict stale ("code changed since —
re-verify") rather than presenting a confident lesson on rewritten code.
- CONTESTED surfaced distinctly (useful N / dead-end M), not averaged away.
- DEAD-ENDS stay QUERY-SCOPED — never a node-level status; they appear only
in the report as question -> nodes.
- Read surfaces (explain / query+MCP / GRAPH_REPORT / graph.html) merge the
overlay at read time, sanitized; un-annotated graphs are byte-identical.
Deferred (logged): letting verdicts influence query/seed traversal — the
recommender feedback-loop / Matthew-effect risk means that needs
propensity correction + exploration, not naive biasing.
Builds on the idea in #1441/#1542 (thanks @TPAteeq).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two cross-platform fixes salvaged from #1502:
- to_graphml: nx.write_graphml raises ValueError on None attribute
values, so a node/edge carrying a null field crashed the export.
Coerce None -> "" for node and edge attributes before writing.
- save-result: add --answer-file as an alternative to --answer so long
or multiline answers can be passed via a file instead of a fragile
inline shell arg (notably Windows/PowerShell quoting). Exactly one of
--answer / --answer-file is required.
The rest of #1502 (a version downgrade and a hand-edited generated
skill-windows.md that fails skillgen --check, plus duplicated
windows-scripts) is left for rework on the PR.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
to_obsidian wrote one note per node straight into the target directory and
unconditionally replaced .obsidian/graph.json. Pointing --obsidian-dir at a real
vault could therefore clobber a user note whose name matched a graph node
(Database.md) and destroy the user's graph-view settings — silently, no backup,
irreversible.
graphify now records the files it owns in .graphify_obsidian_manifest.json and
refuses to overwrite any pre-existing file it didn't create: such a file is skipped
and reported in a single aggregated warning. A re-run still updates graphify's own
notes (they're in the manifest), and .obsidian/graph.json is only written when it
doesn't already exist or graphify owns it. The default graphify-out/obsidian output
and the flat note layout are unchanged.
Added regression tests: existing-vault preserves user note + .obsidian settings,
empty dir still gets the full vault, and a re-run updates own notes but not a
user-added file. The CHANGELOG also records the @oleksii-tumanov Java fixes
(#1512/#1510) and the @nuthalapativarun Windows GBK fix (#1505) committed just prior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
to_canvas sized each community group box for a ceil(sqrt(n))-column grid but the
placement loop hardcoded 3 columns, so any community bigger than ~9 members
rendered as a cramped 3-wide strip in an over-wide, mostly-empty box (and the box
width/height didn't even agree — w used sqrt(n), h used /3). The column count is
now computed once per community (inner_cols) and reused for box width, box height,
and card placement, so the cards fill the box. Cosmetic, no data change.
Ported from PR #1459 by @TPAteeq onto current v8 (clean: only the grid math
changed, the #1457 dedup helper is untouched). Verified the geometry on a real
canvas: n=25 -> 5x5 grid with every card inside its box; n=10 -> 4 columns.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
to_obsidian / to_canvas / to_wiki keyed filename dedup on the exact-case name,
so two labels differing only by case (e.g. `References` vs `references`) counted
as non-colliding and the second write clobbered the first on case-insensitive
filesystems (macOS/APFS, Windows/NTFS) — silently, no suffix, no warning.
Dedup now folds case (keyed on the lowercased name) while emitting the
original-case filename, so any pair that would collide on disk gets a numeric
suffix. The obsidian/canvas dedup is one shared helper (`_dedup_node_filenames`)
so they can't drift; wiki's slug dedup gets the matching fix; the `_COMMUNITY_*`
overview notes (which had no dedup at all) are covered; and a generated `base_1`
is re-checked so it can't overwrite a node literally labelled `base_1`.
Ported from PR #1457 by @TPAteeq onto current v8. Verified with a rigorous
edge-case battery (case-only collision, base_1 literal re-check -> base_1_1,
community-label case fold, determinism) plus the PR's tests; full suite 2404
passed, ruff + skillgen clean.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
An all-punctuation node label (e.g. `@/*` from a tsconfig paths entry) survived
the unsafe-char strip in `to_obsidian`'s `safe_name()` as a bare `@`, producing
`@.md`. That filename is valid on disk but empty once a downstream tool re-slugs
on word chars — qmd's handelize() reduces "@" -> "" and raises, aborting the
entire `qmd update` (every collection on the machine stops reindexing).
Require at least one word char in the stem; otherwise fall back to "unnamed"
(the existing dedup handles collisions). Applied to both safe_name occurrences.
Fixes#1409
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
to_canvas built cards solely by iterating communities, so a graph with no
community data (--no-cluster builds, or a missing analysis sidecar) wrote the
empty 32-byte {"nodes":[],"edges":[]} shell while notes rendered fine. Fall back
to one synthetic community covering every node so the canvas reflects the graph.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
export.py: to_json now accepts community_labels and writes community_name onto
each node. Previously cluster-only wrote labels only to GRAPH_REPORT.md,
graph.html, and .graphify_labels.json — graph.json stored only the numeric cid,
so query/MCP showed blank or numeric community values (#1305).
__main__.py: pass community_labels=labels to to_json in cluster-only path.
explain command now prefers community_name over raw numeric community field.
serve.py: query and get_node read paths prefer community_name over community,
with fallback so old graphs without the field still work. Adds --graph flag as
an alias for the positional argument in graphify-mcp/_main(), fixing
"unrecognized arguments: --graph" for users following the documented pattern
shared by every other graphify subcommand (#1304).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- export.py: guard to_obsidian/to_canvas against dangling community member IDs
(KeyError crash when a node in communities dict is absent from graph, #1236)
- detect.py: NFC-normalize path before hashing Office sidecar filename to fix
macOS NFC/NFD mismatch causing --update to re-extract all Office files (#1226)
- extract.py: add _is_config_json() to skip data JSON files (only extract
package.json, tsconfig.json, eslint, deno, JSON Schema etc.) eliminating
561 orphan key-nodes on large repos (#1224)
- llm.py: add GRAPHIFY_LLM_TEMPERATURE env var + _resolve_temperature() helper;
auto-omit temperature for o1/o3/o4/gpt-5 reasoning models that reject temp=0;
mirrors GRAPHIFY_MAX_OUTPUT_TOKENS precedence pattern (#1191)
- tests: 20 new regression tests across obsidian, detect, extract, llm_backends
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Guards _norm, _norm_label, and _strip_diacritics against None node labels that cause TypeError in unicodedata.normalize(). Fixes#1194. Consistent with existing security.py:270 precedent.
Co-authored-by: freiit <freiit@users.noreply.github.com>
Makes the FalkorDB option a first-class sibling of Neo4j in the agent skill,
not just the export CLI:
- --falkordb / --falkordb-push shorthands documented in core.md + the shared
exports.md reference, so they render into all modular platform skills and
read exactly like --neo4j / --neo4j-push. (The aider/devin monoliths are
diff-frozen vs v8 by skillgen's roundtrip guard, so they are left untouched.)
- README command reference switched to the /graphify ./raw --falkordb-push form.
- Documented URI scheme is now falkordb://localhost:6379; the scheme is only
informational (host/port are parsed out), so redis:// or a bare host:port
remain equivalent. Regenerated skill artifacts + expected/ snapshots.
Three-part fix:
dedup.py: Pass 1 exact-merge now skips nodes with an empty source_file.
Previously all no-source_file nodes with the same label landed in one
bucket and were merged, destroying distinct symbols (third-party deps,
standalone functions) that happened to share a short name.
update.md (skillgen + all 13 host variants): the --update merge now
passes both deleted AND changed files to prune_sources, mirroring what
watch._rebuild_code already does correctly. Old nodes for re-extracted
files are pruned before fresh AST is inserted — no fuzzy reconciliation
needed, no cross-file collapse possible.
export.py: anti-shrink guard message now names fuzzy dedup as a
possible cause (not only "missing chunk files"), and advises a full
rebuild as the safe recovery path.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds FalkorDB as a sibling option to the existing Neo4j sink, selected via
`graphify export falkordb [--push redis://localhost:6379]`.
- New push_to_falkordb() in graphify/export.py mirrors push_to_neo4j; FalkorDB
is OpenCypher-compatible so the MERGE/SET upsert queries are identical.
- export falkordb subcommand wired in graphify/__main__.py (cypher.txt when no
--push, direct push otherwise). Auth is optional; target graph defaults to
"graphify".
- falkordb optional extra in pyproject.toml (and in the all extra).
- Tests: CLI cypher generation (CI-safe) + real-FalkorDB integration tests that
skip when no instance is reachable.
- README extras table + command reference and CHANGELOG updated.
#1118 — prune stale AST nodes on full re-extraction (#1116)
Stamps every AST-extracted node with _origin="ast" in extract(). On a
full rebuild _rebuild_code drops any AST-marked node absent from the
fresh output even when its source file survives, fixing stale symbols.
Backward-compat: marker-less nodes from pre-1118 graphs survive one
cycle then self-heal.
#1110 — stop reading images and PDFs as garbage in headless extract
Images route through per-backend vision payloads (base64/data-URI/bytes
for claude/openai/bedrock); non-vision backends get _strip_pixels for
graceful degradation. PDFs reuse pypdf. 5MB cap, 20-image chunk limit.
#1159 — Salesforce Apex extractor (.cls, .trigger)
Regex-based extractor: classes, interfaces, enums, methods, triggers,
SOQL/DML edges. No new dependency. Dispatched as .cls and .trigger.
#1107 — Azure OpenAI Service backend (--backend azure)
Uses AzureOpenAI SDK client (from existing openai package). Auto-detects
when AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT both set. Uses
max_completion_tokens (not deprecated max_tokens).
#1103 — live PostgreSQL introspection (--postgres DSN)
graphify extract --postgres "postgresql://..." introspects tables, views,
routines, and FK relations via information_schema (SERIALIZABLE READ ONLY).
Credentials sanitized on error. New graphify[postgres] extra (psycopg3).
Union-resolved llm.py conflict: Azure functions + bedrock images= param.
Fixed test_image_vision.py mock to accept timeout= kwarg (our #1112).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
to_obsidian and to_canvas built note filenames from node labels with no
length cap, so a label >=255 bytes crashed write_text with OSError. Add a
shared _cap_filename helper that caps on UTF-8 bytes (not chars, so CJK
labels don't slip past) and appends an 8-char hash of the full label when
truncating, so two distinct labels sharing a long prefix stay distinct.
Both safe_name builders route node, community and canvas filenames through
it; wikilinks stay consistent because they read the same filename dict.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
reconfigure stdout/stderr to UTF-8 at startup; replace → and — in all
print statements with ASCII equivalents as belt-and-suspenders fallback
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
When a graph exceeds the viz node limit, to_html() builds a
community-aggregated meta-graph and recursively calls itself.
The recursive call never carried hyperedges onto the meta-graph,
so graph.html always emitted const hyperedges = [] even when
graph.json contained plenty.
This fix remaps hyperedge node references from semantic node IDs
to community IDs before the recursive call, so hyperedge regions
render correctly in the aggregated view. Hyperedges that collapse
to fewer than 2 distinct communities are dropped (they wouldn't
render as a polygon anyway).
Fixes#1005
* feat(bash): harden extractor — literal filtering, entrypoint nodes, AST-ancestry-aware command detection
Builds on tree-sitter-bash extractor from #866. Two correctness/security
improvements to bash extraction in graphify/extract.py:
1. Reject command/process substitutions at extraction time. Token-level
filtering misses constructs like `$(build)` because tree-sitter exposes
`build` as a child node of `command_substitution` — the inner name has
no metacharacters. Added `is_inside_expansion(node)` that walks
`node.parent` until it finds `command_substitution` or
`process_substitution`. Used as a gate in both `walk` and `walk_calls`.
Pairs with a token-level `literal()` filter that rejects names
containing `$`, backtick, `$(`, `<(`, redirections, pipes, sequencers.
2. Entrypoint node. Every .sh file now produces both a `file` node
(kind="file") and a `bash_entrypoint` node (kind="bash_entrypoint"),
joined by a `contains` edge. A separate top-level `walk_calls(root,
entry_nid, ...)` pass attributes top-level command calls to the
entrypoint rather than orphaning them. Matches the entrypoint pattern
other-language extractors use. Node metadata gains language+kind.
Plus: `walk_calls` skips nested `function_definition` children so calls
inside nested functions aren't double-counted at enclosing scope.
Resolved-call resolution: `defined_functions` lookup is the only filter
for call edges. User-defined functions named like external commands
(install, find, git, ...) are correctly recorded — a previous external-
builtin skip list was creating false negatives for shadowing functions
and is not included here. Skip list belongs with raw/unresolved call
recording (not in this PR).
Devtools (bundled): pyproject.toml gains [dependency-groups] dev (ruff,
pyright, pre-commit, hypothesis, pip-audit) plus minimal [tool.ruff],
[tool.ruff.lint], [tool.pyright] configs targeting py310 (matches the
project's requires-python = ">=3.10").
Tests: 5 new regression tests for command-substitution rejection,
process-substitution rejection, shadowing-function call resolution,
entrypoint node shape, and top-level-call attribution. 826/826 pass
(was 821); 15/15 bash-relevant tests pass (was 10).
* feat(detect): parse macOS/BSD and GNU env(1) shebang option forms
Upstream's _shebang_file_type parses shebangs via line[2:].split() and only
handles `#!/usr/bin/env <interp>`. Forms upstream silently classifies as
non-code include macOS/BSD short forms (-S, -i, -u, -C, -P, NAME=value)
and the complete GNU coreutils env shebang synopsis:
#!/usr/bin/env -[v]S[option]... [name=value]... command [args]...
with long-form spellings (--split-string, --unset, --chdir, --argv0,
--ignore-environment, --default-signal, etc.), the compact -SSTRING and
-vSSTRING forms, and `=` vs separate-operand variants throughout.
Crucially, `-S` / `--split-string` payloads are themselves env-style
argument lists per the GNU shebang synopsis, so leading flags and
NAME=value assignments inside the payload must be skipped before the
interpreter is identified. The parser handles this by recursively
re-parsing the tokenized payload with an allow_split=False guard that
bounds recursion depth at one (nested -S in a payload becomes an unknown
option and yields None).
Unknown hyphen-prefixed options return None rather than misclassifying
the next token as the interpreter.
_shebang_file_type becomes a 4-line wrapper. Read buffer raised 128 -> 256
to accommodate longer env -S strings.
Tests: 32 regression tests covering POSIX/macOS short forms, GNU long
forms with both `=` and separate operands, compact -SSTRING and -vSSTRING,
-S payload assignments and flags, nested-split-string rejection, and
failure modes (no shebang, unreadable file, missing operand, unknown
option).
* fix(skills): enforce semantic fragment validation in OpenCode + Codex merges (#825)
Closes#825. Adds graphify.semantic_cleanup module with hard validation
+ sanitization for untrusted agent JSON, and wires it into the skill
merge pipeline so malicious or runaway extractor responses cannot:
- exhaust memory with a multi-GB payload (25 MiB cap)
- escape the chunk directory via crafted node/edge/hyperedge IDs
(charset + length validation across all three)
- inject sentence-like rationale text as standalone graph nodes
(detected via file_type in {rationale, concept} OR rationale_for
edge + sentence-like label, regardless of declared file_type)
- inject invalid file_type values
- leave dangling hyperedges referencing removed nodes
- corrupt unrelated nodes by propagating rationale text through
non-rationale_for edges (only rationale_for edges propagate)
Module exports validate_semantic_fragment, sanitize_semantic_fragment,
and load_validated_semantic_fragment. Wired into skill-opencode.md and
skill-codex.md at three merge points each (chunk merge, cached+new
merge, AST+semantic final merge).
Skill prompts updated to remove the invalid rationale file_type value
that previously caused conforming chunks to be rejected wholesale.
Valid set is now {code, document, paper, image}.
Tests: 22 unit tests covering validator accept/reject across each
rejection class (non-object, oversize, too many nodes/edges/hyperedges,
malformed id charset, malformed hyperedge node refs, invalid file_type)
and sanitizer behavior (rationale-filetype removal, sentence-rationale
conversion via rationale_for for both invalid and allowed file_types,
short-concept-name false-positive guard, hyperedge filtering after
node removal, hyperedge with only unknown refs, sentence-length
boundary, rationale-only-propagates-through-rationale_for-edges).
880/880 tests pass.
* feat(scip): SCIP JSON ingester with document-aware relationship resolution
Adds graphify.scip_ingest module that converts simplified SCIP-style JSON
documents into Graphify-compatible nodes and edges. Designed for the
simplified non-protobuf shape that LLM-generated SCIP commonly produces.
Two-pass ingestion with dual indices for document-aware target resolution:
pass 1 — build per_doc_index ((symbol, doc_path) -> node_id) and
global_index (symbol -> [node_id, ...]) across every valid
symbol in every valid document. Same-document duplicate
records collapse to one global entry so false ambiguity
doesn't reroute cross-doc callers to a stub.
pass 2 — emit nodes for indexed symbols, then walk relationships.
Resolution order:
1. same-doc match (per_doc_index)
2. unique cross-doc match (global_index[symbol] len == 1)
3. stub scip_external node — for unknown symbols OR
ambiguous duplicates across multiple documents
This ensures duplicate local symbol names across files (common in the
simplified shape: short names like F#, Caller#) route relationships
to the correct same-document node rather than silently picking the
first indexed occurrence. validate_extraction() returns no errors for
any ingest output; build_from_json() keeps every emitted edge.
Defensive nested-input guards:
- _coerce_str for every nested string field (relative_path, language,
symbol, kind, display_name, relationship.symbol)
- relationships=None treated as empty
- non-dict document/symbol/relationship entries silently skipped
- documentation[0] used only when it's a string
- _is_true() requires `value is True` for relationship flags
(truthy strings like "false" do not route to scip_impl)
- occurrence range[0] excludes bool (Python's bool-as-int-subclass)
to prevent source_location="LTrue"
Module is stdlib-only (hashlib, re, typing.Any). Not wired to the CLI
in this phase — importable as `from graphify.scip_ingest import
ingest_scip_json`.
Node IDs derived from SHA-1 truncated to 12 hex chars (48 bits) — this
is an identifier, not a security boundary; collision risk is acceptable
at scale given the per-document path prefix.
Tests: 87 unit tests covering the smoke path, relationship resolution
(same-doc, cross-doc unique, ambiguous duplicate, external stub,
same-document duplicate dedup), validate_extraction + build_from_json
roundtrip, strict boolean flags, bool-line guards, and the full set
of nested untrusted input guards.
1044/1044 tests pass.
* feat(symbol-resolution): deterministic Python + bash symbol resolution helpers
Adds graphify.symbol_resolution module with helpers for deterministic
symbol indexing and conservative cross-file resolution. Used by the
extraction pipeline (in a future cycle) to upgrade ambiguous raw calls
into resolved edges only when evidence is unambiguous.
Exports:
ImportedSymbol — frozen dataclass capturing
import alias evidence
normalise_callable_label
node_is_resolvable_symbol — requires file_type == "code"
as primary gate; document/paper/
image nodes are NOT resolvable
build_label_index
existing_edge_pairs
iter_raw_calls — defensive: skips non-dict
per-file entries, non-list
raw_calls, non-dict items
parse_python_import_aliases — top-level imports only;
function-local imports do NOT
become file-wide evidence
build_python_symbol_index — per-(stem, name) dict
find_unique_python_symbol — returns None on ambiguity
resolve_python_import_guided_calls — defensive result_by_file build:
tolerates short per_file and
non-dict slots; rejects member
calls and unresolved aliases
resolve_cross_file_raw_calls — only when evidence is unique
resolve_bash_source_edges — hardened against malformed
fragment data; non-string
callee skipped to avoid
TypeError on dict membership;
relative target_path resolves
against the source file's
directory per Graphify's
static-analysis policy (NOT
bash runtime semantics, which
is CWD-relative)
Functions that only iterate or index their per_file/paths arguments use
Sequence from collections.abc for proper covariance. Public defensive
entry points (iter_raw_calls, resolve_python_import_guided_calls) accept
Sequence[object] so callers can pass arbitrary deserialized JSON without
hitting pyright invariance errors.
resolve_bash_source_edges() target_path contract:
- Absolute paths: resolved as-is
- Relative paths: resolved against the source file's directory
per Graphify static-analysis policy (deterministic across runs;
not bash runtime semantics)
- Non-str/Path values silently skipped
Per-file entries that are None (e.g. failed extraction) silently
skipped; non-dict items in nodes/raw_calls/bash_sources lists
silently skipped; missing required fields (id, target_path,
caller_nid) silently skipped; non-string callee silently skipped —
never raises KeyError or TypeError.
Module is stdlib-only (ast, re, dataclasses, pathlib, typing,
collections.abc). Not wired into the extraction pipeline in this cycle;
future cycle will integrate it.
Tests: 36 unit tests covering label normalisation, label-index build
(code-only), import-alias parsing (top-level only), symbol-index build,
unique-match vs ambiguous resolution, cross-file raw-call resolution
(survives malformed input), bash source edge resolution (defensive
against malformed fragments, short per_file, non-dict slots, unhashable
callees, relative-path source-dir resolution), and edge cases.
* feat(security): cap graph.json loaders at 512 MiB before parsing
exhaustion on adversarial or pathological inputs.
- graphify.security: add _MAX_GRAPH_FILE_BYTES + check_graph_file_size_cap
- graphify.serve._load_graph: call cap after existence check
- graphify.__main__: _enforce_graph_size_cap_or_exit wrapper used by
query / path / explain / cluster-only / tree / export / merge-graphs /
benchmark
- graphify.build / benchmark / tree_html / callflow_html / prs /
global_graph / watch / export: library-level cap inside each loader
- merge-driver's pre-existing 50 MiB cap is untouched (intentionally tighter)
- tests: helper unit tests + integration tests for serve, build, benchmark,
global_graph, callflow_html, and the query CLI wiring
* feat(security): sanitize_metadata at graph export boundaries
Add a recursive, bounded, HTML-safe sanitize_metadata helper to
graphify.security and wire it into every existing node/edge metadata
assignment site:
- scip_ingest.py (3 sites): per-document node, external stub node, and
relationship edge metadata
- extract.py (1 site): bash extractor's add_node metadata
- symbol_resolution.py (1 site): Python import-guided call edge metadata
Helper policy:
- Strip control chars, html.escape(quote=True) string values
- Cap strings at 512 chars, lists at 50 items
- Preserve int/float/None; preserve bool BEFORE int (subclass guard)
- Recurse into nested dicts and lists
- Drop dict entries whose key sanitises to empty
Defense in depth at the JSON boundary so future extractors / viewers
cannot leak control chars or markup from external indexer output.
* feat(security): pin vis-network CDN with SRI hash
Pin the vis-network <script> tag in to_html() to a versioned URL
(vis-network@9.1.6) with a sha384 Subresource Integrity hash and
crossorigin="anonymous". Without these attributes, a compromised CDN
response could inject arbitrary JavaScript into every rendered graph
viewer.
Hash verified live against
https://unpkg.com/vis-network@9.1.6/standalone/umd/vis-network.min.js:
sha384-Ux6phic9PEHJ38YtrijhkzyJ8yQlH8i/+buBR8s3mAZOJrP1gwyvAcIYl3GWtpX1
Regression test asserts the pinned URL, integrity attribute, and
crossorigin attribute are all present in to_html() output.
Follow-up: tree_html.py (D3) and callflow_html.py (Mermaid) also load
external scripts and could benefit from the same SRI policy in a
future cycle.
* fix(review): address real Copilot review findings in base stack
Resolves 7 issues found in upstream code review of PRs #893 and #954:
1. extract.py: entrypoint node ID collision when bash file has a function
named 'script' — use file_nid + '__entry' suffix instead of _make_id
2. extract.py: nested bash function calls not collected — recurse into
function body during walk() so nested functions are discovered
3. extract.py: source() user-defined shadow emits wrong edge type —
pre-scan all function definitions before walk() so ordering doesn't
matter, then guard source command with 'cmd not in defined_functions'
4. extract.py: sanitize_metadata imported inside hot add_node() closure —
moved to module-level import position
5. symbol_resolution.py: _bash_make_id() diverged from extract._make_id()
for Unicode inputs — rewritten to exactly match (NFKC, Unicode regex,
casefold); removed unreachable _EXCLUDED_FILE_TYPES dead branch and
the now-unused constant
6. semantic_cleanup.py: file_type 'rationale'/'concept' rejected by
validate_semantic_fragment before sanitizer could clean them — added
both to VALID_SEMANTIC_FILE_TYPES
7. scip_ingest.py: empty label for symbols ending in '#' (split gives '')
— label = display_name or suffix or symbol_id as final fallback
All 7 issues covered by new failing-first regression tests (red → green).
Full pytest suite: 1239 passed, 4 pre-existing env-specific failures.
* fix(review): address PR #956 Copilot findings in watch.py and symbol_resolution.py
- watch.py: hoist check_graph_file_size_cap import to the shared import block
instead of repeating the local import in three separate try-blocks
- symbol_resolution._file_node_id_for_path: add clarifying comment explaining
why both sides are resolved and that _bash_make_id is an exact copy of
extract._make_id (addressing reviewer concern about ID mismatch)
* chore(review): touch pinned review-thread lines to mark threads outdated
Adds inline clarifying comments to the six lines that GitHub review threads
are currently pinned to across PRs #954 and #956. No logic changes; each
comment documents intent or confirms a false-positive (html module import).
* feat(diagnostics): report multigraph edge-collapse risk
Add graphify.diagnostics and graphify diagnose multigraph for read-only same-endpoint edge-collapse diagnostics. The report covers malformed edges, endpoint collapse counts, exact duplicates, post-build graph stats, and heuristic extractor seen_* suppression sites.
Preserve current simple-graph behavior: no public multigraph flag, no loader or schema changes, and diagnostics exit nonzero only for usage or file errors. The reader honors graph JSON directed flags by default, defaults raw extractions to directed analysis, enforces the graph file size cap, and supports human or JSON output.
* feat(multigraph): add runtime compatibility probe
New module graphify.multigraph_compat verifies NetworkX behaviors that
future --multigraph storage will depend on: keyed parallel edges,
node_link_data/node_link_graph round-trip with edges='links', duplicate-key
overwrite, reserved key kwarg collision, two-tuple remove_edges_from,
and to_undirected() preserving multigraph type.
Behavior probe, not version check. Both NX 3.4.2 (Py 3.10 lane) and
NX 3.6.1+ (Py 3.11+ lane) pass. Result cached for the process lifetime.
No call sites added — this PR adds the API surface only. Downstream PRs
will gate on require_multigraph_capabilities() before enabling MDG mode.
Refs: Wave 1 MultiDiGraph implementation order.
* test: filter known third-party analyze warnings
---------
Co-authored-by: vampyre <vampyre@local.net>
#796: add edge_data()/edge_datas() helpers in build.py that tolerate
MultiGraph/MultiDiGraph; replace all G.edges[u,v] 2-tuple call sites in
__main__.py, serve.py, wiki.py, export.py, analyze.py, benchmark.py;
fix same pattern in 10 skill file inline heredocs
#795: all 12 skill files now short-circuit on /graphify --help or -h
and print the Usage block without running any pipeline steps
#792 (hollow response): add _response_is_hollow() predicate in llm.py;
when Ollama (or any backend) returns empty/null/whitespace content or a
parsed result with no nodes/edges, rewrite finish_reason="length" so
_extract_with_adaptive_retry bisects the chunk instead of silently
dropping it; applied to _call_openai_compat, _call_claude, _call_bedrock
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Adds support for Quarto markdown (.qmd) files by:
- Adding '.qmd' to document file extensions in detection
- Updating export logic to handle .qmd in filename sanitization
- Adding .qmd extractor dispatch using the existing markdown extractor
- Updating watch comments to include .qmd files
- Add Fortran support (26th language): .f/.F/.f90/.F90/.f95/.F95/.f03/.F03/.f08/.F08
via tree-sitter-fortran; capital-F files preprocessed with cpp -w -P
- Add graphify export {html,obsidian,wiki,svg,graphml,neo4j} CLI subcommands
- Add graphify query/path/explain CLI subcommands
- Reduce skill.md from 63KB to 47KB by replacing Python heredocs with CLI calls
- Extend to_html() with node_limit param for auto-aggregation on large graphs
- Add integration tests for all export/query/path/explain subcommands
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>