Status: v0.2.x is released and under active development. The CLI, Neo4j connector, core engine, built-in core and PII packs, deterministic sampling, baseline lookup, and report writers are available. The public contracts are frozen for v0.
What GraphCheck can check
GraphCheck supports three kinds of checks in the same suite:- Conformance checks apply reusable built-in rules to graph structure and properties.
- Competency checks run an authored read-only Cypher query and assert its rows, columns, uniqueness, emptiness, or exact/contained results.
- Drift checks compare current node counts, relationship counts, or property coverage with a stored baseline.
PII checks are explicitly heuristic and sampled. Reports include confidence metadata and graph
locations, but never include matched property values or claim complete PII discovery.
Requirements
- Python 3.12 or 3.13
- Neo4j Python driver 5.20 through 6.x
- Neo4j Server 5.26 LTS or a tested calendar-version release
- Cypher 5, or Cypher 25 on the tested calendar-version server
uvonly when using the development workflow shown below
graphcheck init and graphcheck debug. A missing optional
capability blocks only checks that declare it; those checks are recorded as unsupported instead of
silently passing.
Install
Install the core GraphCheck CLI from PyPI:generate add-on to use graphcheck generate and its supported model providers:
mcp add-on to expose GraphCheck through graphcheck mcp serve:
Development install
Clone the repository and create its locked development environment:graphcheck command. Prefix them with uv run when working
from the development environment.
Quickstart
1. Create a project
Run the initializer in the directory that should contain your checks:graphcheck.yml.
2. Configure Neo4j
First check the edition of the running DBMS:
For Enterprise/Developer, an administrator assigns Neo4j’s built-in
reader role to the account
used by GraphCheck. That account must have no other assigned role except the automatic PUBLIC
role. GraphCheck therefore rejects admin, architect, publisher, editor, and custom roles.
Edit profiles.yml. This Enterprise/Developer example uses an account assigned the built-in
reader role:
password_env overrides a literal password when the variable is set. If the environment variable
is absent, GraphCheck falls back to password when one is configured. For the quickest local
setup, edit the generated inline password value. For CI or shared environments, remove the inline
value, keep password_env, and export that variable in the process that runs GraphCheck.
On Neo4j Enterprise/Developer, the account must have only the built-in reader role and the
automatic PUBLIC role. During init, debug, and the CLI run preflight, GraphCheck reads the
current user’s roles with SHOW CURRENT USER. A missing reader role or any additional role is
rejected as neo4j.credential_not_read_only. If Enterprise cannot return the roles, GraphCheck
fails closed as neo4j.credential_read_only_unverified.
Neo4j Community has no roles and gives every user implied administrator privileges, so it cannot
provide a server-enforced read-only credential. GraphCheck explicitly supports Community by
skipping the unavailable Enterprise RBAC gate. In both editions, every customer-authored query is
separately planned with EXPLAIN; GraphCheck executes only Neo4j query type r, so a write-capable
query is rejected without modifying the graph. Driver read routing alone is not an authorization
boundary.
In Desktop 2, open the instance’s Query tool as the existing neo4j administrator. On Neo4j
Enterprise/Developer, grant the built-in role to the user configured in profiles.yml:
admin, architect, publisher, editor, or custom role from that user; reader and
the automatic PUBLIC role must be its complete role set. GraphCheck no longer requires custom
roles or individually granted privileges.
Set the matching password in the same shell that starts GraphCheck:
user: neo4j (or another Community user) and configure
its password normally. Community users are admin-equivalent by design, so GraphCheck skips the
unavailable RBAC audit and relies on its planner guard described above.
Use bolt://host:7687 for a direct non-TLS local server. Use neo4j+s://host:7687 for routing with
CA-validated TLS (including the URI supplied by Aura), or neo4j+ssc://host:7687 only when the
deployment intentionally uses a self-signed certificate. The URI scheme must match the server.
Verify connectivity, server metadata, visibility, graph counts, and check capability requirements:
Fix:;
graphcheck run preserves the same diagnostic in results.json and the HTML report.
3. Add a first suite
Replacechecks/example.yml with a baseline-free suite that matches labels in your graph:
4. Run the checks
--suite and --select are repeatable. Repeated tag selectors use OR semantics. --fail-fast
stops after the first error-severity failure or error, while retaining later selected checks as
explicitly not run.
Every prepared run writes:
results.json follows a versioned, stable contract; the file records its own schema_version. report.html embeds its styling and
interaction script and has no external assets or network calls, so it can be opened and shared
offline. The run-id directory preserves history; latest is a consistently published convenience
copy of the newest run.
Use graphcheck run --redact (--redacted remains an alias) when the generated artifacts will be
shared. Mask mode preserves
verdicts, scores, run-level counts, keys, and container structure while replacing query text,
parameter, expected, and measured literals; check names and provenance; partial reasons; diagnostic
messages and fixes; source hashes and target identifiers; and evidence messages/element values with
[REDACTED]. Suite, check, and tag identifiers receive consistent ordered aliases so their
relationships remain intact. Redacted artifacts use a target-neutral redacted_<timestamp> run ID.
Redaction also compares the final artifact with its collected source literals, allowing collisions
only in explicitly safe structural fields such as timestamps, versions, enums, and error codes.
Every mask-mode JSON and HTML write verifies the mask, alias, and neutral-ID policy before export.
Redacted HTML reports omit all target metadata and graph counts. Check cards show the pattern under
the check name and omit the details/evidence toggle and its Expected, Measured, and Compiled Cypher
sections.
To create a safe sidecar from an existing run:
--output, the command writes results.redacted.json beside the source and never
overwrites the original.
Exit codes
GraphCheck uses stable CI-oriented exit semantics:
Exit
2 deliberately distinguishes incomplete or warning-level outcomes from both success and a
hard failure.
Drift baselines
Drift checks supportnode_count, relationship_count, and property_coverage. The CLI resolves
baseline JSON from <artifacts>/baselines/, which is .graphcheck/baselines/ by default:
baseline: release-2026-07resolvesrelease-2026-07.json.baseline: latestselects the lexicographically newest.jsonfilename.
graphcheck profile captures timestamped baselines in that directory. A missing, invalid, or
incomplete requested measurement is an explicit check error, never a pass. A partial baseline
can still serve a drift check as long as it contains a present, valid value for the measurement
that check requests.
Generate check suggestions
This feature requires the separately installedgenerate add-on:
graphcheck generate turns the latest baseline profile and optional, explicitly named domain
documents into non-deterministic check suggestions. The command discloses the destination and exact
data categories before calling the configured provider. It never sends graph records, property
values, credentials, target metadata, fingerprints, or profiler failure text.
Generation is opt-in. Add one of these strict blocks to graphcheck.yml:
api_key_env.
Google’s Gemini structured-output models
and hosted Gemma models use the
Gemini API subject to Google’s current quotas and
data terms. Models named gemini-* use native structured output and the full GraphCheck proposal
contract. Other Google model names retain the conservative Gemma tool path: conformance targets
completeness/uniqueness, competency expectations target returned columns, and drift targets
labeled node counts. Gemma publishes a valid partial first batch without a slower correction, so
written may be below requested. Ollama requires an explicit base_url and may omit the key.
Then run:
generated: true at both file and check level, so the engine
validates but skips them without querying Neo4j. Review identifiers, Cypher, expectations,
thresholds, and cost before removing both applicable markers to activate a check.
Reliability and safety
- Neo4j execution is read-only and fails closed unless the planner classifies a statement as read.
- The Enterprise connection preflight fails unless Neo4j reports only the built-in
readerrole andPUBLIC, or if GraphCheck cannot inspect the current user’s roles. - Built-in Cypher keeps labels, relationship types, property names, regexes, thresholds, and values in parameters rather than interpolating user data into query text.
- One broken query or evaluator error is isolated to its check unless fail-fast or the run deadline stops later work.
- Fail and warning verdicts require graph-element or aggregate-scope evidence.
- Sampled checks use deterministic per-check selection and report a 95% Wilson confidence interval; exhaustive runs are labeled exact.
- Missing labels, relationship types, or explicitly selected properties are errors rather than empty passes.
- Result JSON is revalidated against its typed contract before it is written.
Optional anonymous telemetry
Telemetry is off by default. GraphCheck does not create a telemetry client or send events until a user explicitly opts in. Persistent consent usesgraphcheck telemetry enable and is revoked
with graphcheck telemetry disable. The process-only GRAPHCHECK_TELEMETRY=1 environment
override opts in a single process and takes precedence over stored consent; unset it or set it
to 0 to stop it.
When enabled and delivery is configured, GraphCheck sends only structural and aggregate signals
such as command/run occurrence, timings, execution counts, operational outcomes, and coarse
runtime information. It never sends queries, graph schema names or values, database or project
identity, credentials, check identities, check results or verdicts, command arguments, paths, or
free-form errors. Delivery is asynchronous and best-effort, so telemetry failures never change CLI
behavior. See the telemetry disclosure for the complete event and field
inventory.
Configuration reference
graphcheck.yml has three required strict fields plus optional concurrency and generate
settings:
checks and artifacts paths are resolved from the project root. Suite discovery is
recursive and includes .yml and .yaml files. Every discovered suite is read and validated
directly on each command before suite-id filtering; GraphCheck does not create a suite-discovery
cache file. concurrency is a positive worker limit; the default is 1, and
graphcheck run --concurrency N overrides the project value.
For agent integration, result consumption, and programmatic check authoring, see the
GraphCheck agent guide.
Development
Install the locked development environment and run the same checks as CI:GRAPHCHECK_NEO4J_INTEGRATION=1; the testcontainers
suite covers the connector and engine against supported Neo4j versions. The customer-scale
performance test is separately opt-in and requires a preloaded graph of at least 10 million nodes.
The repeatable performance gates and measurement helpers live under
tests/performance.
See CONTRIBUTING.md for the branch workflow, definition of done, and decision
rights.
