Skip to content

Verification and quality gates

Status: active quality contract
Applies to: local development, CI, differential infrastructure, and releases

Verification answers three separate questions:

  1. Does Oxiland implement its documented Rust API correctly?
  2. Does it match the applicable RDF/SPARQL standards?
  3. Does it reproduce the Redland behavior claimed by the compatibility plan?

A passing test in one layer does not substitute for the others.

Test layers

Layer Purpose Location
Unit conversions, configuration, invariants, errors module-local tests
Rust integration public workflows and feature combinations tests/
Conformance RDF and SPARQL standards compatibility/conformance/
Differential Oxiland versus native Redland compatibility/fixtures/
C contract (0.9 preview) headers, symbols, allocation, callbacks crates/oxiland-capi/tests/
Python package wheels, typing, pytest python/
Downstream real language bindings and applications CI-managed manifests
Fuzz/property (planned expansion) malformed inputs and lifecycle sequences fuzz/

Planned locations are created with the milestone that first needs them.

Fixture requirements

Every compatibility fixture has:

  • a stable ID matching one or more inventory entries;
  • setup and cleanup that are isolated and deterministic;
  • backend operations expressed without implementation-specific shortcuts;
  • expected result category and normalization profile;
  • assertions for both success and relevant failure paths;
  • source/oracle version metadata.

Tests involving unordered results must not rely on incidental iteration order. Network-dependent behavior is captured in hermetic fixtures or runs in a separate explicitly non-hermetic suite.

Differential harness

Fixtures should be data-driven so one case can execute through both backends. Each case contains setup, operations, expected result class, and normalization rules. Backend runners emit a common JSON result containing values, errors, logs, and relevant state.

Differences are classified as:

  • implementation defect;
  • expected non-semantic formatting variation;
  • undocumented Redland behavior that becomes part of the compatibility target;
  • accepted deviation with a published rationale.

Golden outputs alone are insufficient because they can encode a mistaken interpretation. Where possible, the native Redland runner is the oracle.

The harness produces machine-readable output and a human-readable diff. A fixture passes only if both runners complete, normalization succeeds, and the result comparison is equal. Crashes, timeouts, skipped runners, and missing oracle metadata are not passes.

Standards conformance

Relevant W3C manifests should be pinned and run through Oxiland's public API, not only inherited from Oxigraph's upstream test claims. Expected upstream deviations are recorded with version and issue links.

Conformance categories include:

  • RDF term construction and equality;
  • Turtle, TriG, N-Triples, N-Quads, and RDF/XML;
  • SPARQL query, update, protocol-independent dataset semantics, and results;
  • RDF 1.2/SPARQL 1.2 only when the corresponding Oxiland feature is promised.

Safety and robustness

  • Miri covers safe abstractions where practical.
  • AddressSanitizer and LeakSanitizer cover the C ABI and native harness.
  • UndefinedBehaviorSanitizer covers C shims where supported.
  • Fuzzers retain a checked-in regression corpus.
  • Panics are acceptable only for documented programmer invariants; malformed RDF, queries, files, options, or C inputs must not panic.
  • Thread and callback tests include re-entry and concurrent destruction where the API permits them.

CI matrix

Required on every PR and on main/release:

  • stable Rust checks (fmt, Clippy, tests, docs, examples, inventory, documentation links, public-API snapshot);
  • safe Rust workspace tests on Linux, macOS, and Windows;
  • dedicated Fjall persistence tests;
  • Rust 1.87 MSRV Clippy and tests;
  • RustSec audits for the workspace and Python extension lockfiles;
  • non-breaking semver compatibility against 0.9.0 (forced to minor-release policy for the pre-1.0 API) and packaged-crate verification;
  • Python pytest, Pyright, and runnable examples;
  • Linux, macOS, and Windows wheel builds for CPython 3.10–3.14;
  • metadata, license, typing, native-extension, and CycloneDX SBOM validation for every wheel;
  • clean install/import smoke tests for every produced wheel;
  • RDF I/O conformance and compatibility harness smokes.

Planned broader coverage includes:

  • sanitizer-enabled Linux/macOS C ABI tests;
  • expanded license policy and generated dependency inventory checks.

Workflow permissions default to read-only, third-party Actions use immutable full commit SHAs, and dependency updates arrive as grouped Dependabot pull requests. Release jobs elevate only the individual permissions needed for OIDC attestations or GitHub release assets. The Rust workspace, independent Python extension, and fuzz workspace each commit their Cargo.lock, making --locked builds and RustSec results reproducible on clean runners.

Security advisories are blocking on tip CI (the same reusable workflow release runs). Tip also validates package-version alignment, cargo publish --dry-run for the library crate, and the full 15-wheel release matrix after per-OS install smokes. PyO3 is at 0.29.0. Oxigraph 0.5.9 still constrains quick-xml to 0.37 (RUSTSEC-2026-0194 / RUSTSEC-2026-0195); tip CI allows only those two IDs, and scripts/check-security-exceptions.py fails as soon as the Oxigraph/quick-xml graph changes so the waiver must be re-reviewed or removed. See R-020 in the risk register.

Nightly or scheduled coverage includes:

  • fuzzing and Miri;
  • full W3C and differential suites;
  • big-endian or cross-architecture checks when infrastructure permits;
  • persistent-store crash and concurrency scenarios;
  • downstream rebuilds and performance baselines.

Persistent storage tests use isolated temporary directories and verify reopen, rollback, interrupted writes, and concurrent access behavior.

Local gates

Before review:

cargo fmt --all --check
cargo clippy --workspace --all-targets --all-features --locked -- -D warnings
cargo test --workspace --all-features --locked
cargo doc --workspace --all-features --no-deps --locked
python3 scripts/check-inventory.py
python3 scripts/check-docs.py
scripts/generate-public-api.sh check

For Python changes:

cd python
python -m pip install --requirement requirements-ci.txt
maturin develop --locked
pytest -q
pyright
python examples/quick_start.py
python examples/select.py
python examples/parse_serialize.py
python examples/persistent.py
maturin build --release --locked

For documentation changes:

python3 scripts/check-docs.py
python3 -m pip install --requirement docs/requirements.txt
python3 -m mkdocs build --strict

Milestones may add native or long-running commands to this baseline. Release artifact checks must operate on the built crate, CLI, or wheel rather than only the workspace source tree. Python release artifacts are the exact CI-built and install-smoked wheels; release jobs do not rebuild them. Their metadata, bundled licenses, PEP 561 files, native extension, and CycloneDX SBOM are checked before provenance attestation and publication.

Release gates by phase

Every 0.x release requires:

  • all local gates on required platforms;
  • updated parity and roadmap status;
  • updated dependency audit and minimum Rust check;
  • successful packaging and clean-install smoke tests;
  • no unexplained regression in the milestone's differential suite.

Additional phase gates:

Starting version Added release blocker
0.2 applicable RDF syntax conformance
0.3 SPARQL query/update facade conformance and smoke harness (Rasqal differential expands later)
0.4 persistence, transaction, and reopen matrix
0.6 complete safe-API inventory and public-API snapshot
0.7 Python wheels, type checks, and pytest matrix
0.8 exported symbols, C examples, and sanitizers
0.9 selected downstream C consumers
0.10 frozen contracts, candidate inventory/C surface, qualification scaffolding, performance candidate, and soak
0.11 demonstrated full Redland 1.0.17 parity, native cross-platform differentials, source compatibility, and binary ABI interchange
0.12 competitive-parity performance gate (ADR-028) on parity-qualified artifacts, with resource budgets and no required-case waiver
0.13 suite-wide faster-than-Redland gate (ADR-029): three independent corrected-runner runs × Linux/macOS/Windows

The 0.11 parity row is a hard gate, not a documentation claim or a waivable target. The machine-generated report must derive every state from raw native Redland and Oxiland executions on each declared target/profile. It covers the complete public denominator, unchanged-source C builds, Redland-built binary interchange, and every applicable observable behavior. It reports numerator, denominator, skips, platform/profile, exact tested revision, and hashes for all inputs and artifacts. Any unreviewed, mapped-only, implemented-but-unverified, or excluded item; missing oracle result; differential mismatch; accepted deviation; quarantine; capability-error substitute; migration-only workaround; stale revision; synthetic pass; or copied profile result blocks 0.11. Safe-Rust-only ownership mechanics may be not-applicable, but their C lifecycle behavior must still pass. The normative definition is in Compatibility.

The 0.10 checked-in bundle remains useful for regression and tooling tests, but it cannot satisfy the 0.11 gate. A 0.11 generator may summarize raw results; it may not turn an allowlist or local smoke test into differential evidence.

Flaky tests are quarantined only with an owner, issue, expiry milestone, and a replacement signal. Quarantined compatibility tests do not count as passing.

Performance verification

Compatibility is primary, but accidental performance cliffs can make a compatible API unusable. Benchmarks track:

  • triple insert/remove and pattern scans;
  • parsing and serialization throughput;
  • query latency and result streaming;
  • persistent reopen and bulk load;
  • peak memory on large streams;
  • C-call and callback overhead.

Budgets and the required comparison suite are established from representative workloads before 0.10 qualification. Benchmark noise is controlled by the protocol below rather than waived.

0.10 performance scaffold

The 0.10 release freezes the faster-than-Redland protocol and ships synthetic candidate fixtures that exercise its fail-closed statistical validator. It does not claim that Oxiland 0.10 was natively benchmarked against Redland on each declared target. The exact parity-qualified artifacts must later pass the protocol before a faster-than-Redland or 1.0-readiness claim is made.

Under that protocol, throughput cases require an Oxiland/Redland median ratio of at least 1.05; latency cases require an Oxiland/Redland median ratio of at most 0.95. In both cases, the 95% bootstrap confidence interval must exclude parity (1.0) on the winning side. A tie, statistically inconclusive result, or loss blocks the eventual performance claim; a geometric mean or win elsewhere cannot hide it.

The comparison protocol is part of the gate:

  • Oxiland's C compatibility surface and Redland's public C API run the same completed, behaviorally equivalent workload with identical inputs and result validation;
  • both use pinned production/release builds, equivalent compiler optimization, the same machine, OS image, allocator policy, storage medium, cache state, and thread limits; Oxiland must be Cargo --release (not debug/dev), and the C harness must not mix a release library with an unoptimized wrapper;
  • cold and warm cases are separated, execution order alternates or is randomized, warm-up is recorded, and enough repetitions are retained to publish medians, dispersion, confidence intervals, and raw samples;
  • the frozen suite covers triple mutation and scans, parse/serialize, query and streaming, persistent reopen and bulk load, and C-call/callback overhead at small and representative large sizes; and
  • peak memory and disk amplification remain within their independently published safety budgets, so speed cannot be bought by an unbounded resource regression.

The suite, datasets, queries, target/profile matrix, toolchain, Redland build, measurement method, and pass thresholds freeze before qualification begins. Required cases cannot be deleted, renamed optional, or waived after a failure. The 0.11 candidate-bound qualification runs the full matrix on controlled benchmark hosts and publishes signed machine-readable and human-readable results. Those 0.11 samples are diagnostic evidence; closing every required win is the 0.12 milestone.

0.12 performance optimization

Milestone 0.12 owns the performance gate on parity-qualified artifacts. Under ADR-028, matched production builds freeze a competitive-parity rule:

  • throughput: Oxiland/Redland median ≥ 0.90, and 95% bootstrap CI lower bound > 0.85;
  • latency: Oxiland/Redland median ≤ 1.20, and 95% bootstrap CI upper bound < 1.40.
  • At least 40 independent samples per case. A later ADR (ADR-029) restored the stricter faster-than-Redland margin when matched evidence sustained it. 0.12 alone does not authorize a blanket “faster than Redland” marketing claim; 0.13 does after nine green cells.

Production compile contract

Rust performance is a function of how the code is compiled. 0.12 evidence is valid only when Oxiland is measured from a production / release build:

  • Cargo profile release via cargo build … --release (qualification also passes --locked);
  • artifacts from target/release (or the equivalent release output directory), never target/debug;
  • debug_assertions disabled (Cargo release default);
  • no RUSTFLAGS / profile overrides that force opt-level = 0 or a dev profile for the measured library;
  • C perf_bench wrappers and Redland built with equivalent native optimization (at least -O2), recorded alongside the Oxiland provenance;
  • flamegraphs and attribution runs use the same release binaries as the timed samples.

The frozen suite sets protocol.require_production_compile: true. scripts/check-performance-gate.py rejects missing build provenance, non- release Cargo profiles, debug artifact paths, and -O0 Redland flags when that flag is set. A fast debug build cannot satisfy the gate.

Additional 0.12 rules:

  • Optimizations must retain a green 0.11 parity checker on the candidate ancestry; behavioral shortcuts are out of scope.
  • Peak RSS and disk amplification must meet independently frozen budgets so speed cannot be bought by unbounded resource growth.
  • Published claims are per-case median ratios with confidence intervals, host, profile, suite revision, and compile provenance—not a geometric mean across the suite.
  • The 0.12 release checker fails closed on synthetic samples, debug/dev compiles, missing resource checks, stale revisions, and incomplete profile coverage.

0.13 suite-wide faster-than-Redland

Milestone 0.13 restores the historical stricter margin under ADR-029:

  • throughput: Oxiland/Redland median ≥ 1.05, and 95% bootstrap CI lower bound > 1.0;
  • latency: Oxiland/Redland median ≤ 0.95, and 95% bootstrap CI upper bound < 1.0;
  • 100 paired AB/BA samples; RSS budgets 1.25;
  • three independent runs per target (Linux x86-64, macOS Apple Silicon, Windows x86-64).

scripts/check-0.13-release.py fails closed on missing cells, duplicate execution_ids, non-production compiles, harness checksum mismatch, and stale revisions. Evidence lives under compatibility/qualification/performance/0.13/.

Metrics

Track separately:

  • inventory items reviewed, implemented, and verified;
  • tests by Redland subsystem;
  • standard conformance pass rates;
  • known behavioral deviations;
  • downstream projects passing;
  • required performance cases beating Redland, with ratios and confidence intervals;
  • parser and FFI fuzzing time without findings.

A single percentage must not combine these categories. Doing so would obscure whether apparent progress represents documentation, implementation, or actual behavioral verification.

Each release publishes numerator, denominator, skipped count, and the exact inventory/suite revision for every percentage.