SDB-26
Open benchmark · CC BY 4.0 GitHub STANDARD.md Plica
Synthetic Document Benchmark · v1.0 public draft

A benchmark for document authenticity

Not marketing accuracy — measurable bypass rate, confidence gaps, and auditable decision artifacts. P · L · I · C · A methodology formalised for regulated file workflows.

01 Overview

SDB-26 measures whether verification systems withstand real synthetic-document attacks.

SDB-26 defines reproducible evaluation for synthetic, edited, and screen-recaptured artifacts in operational conditions — transparent metrics and schema-valid outputs your MLRO can cite, not a vendor slide.

Canonical spec: STANDARD.md · Methodology: METHODOLOGY.md

02 Measurement grid

What we measure

SDB-26 is built around comparable, defender-oriented outcomes:

Metric Meaning Why it matters
BR (Bypass Rate) Share of fraudulent/synthetic documents incorrectly approved Core indicator of control failure
CG (Confidence Gap) Mean confidence on wrongly approved cases Detects overconfident error patterns
GS (Generator Sensitivity) BR segmented by generator/model family Shows where systems break first
FPR (False Positive Rate) Share of genuine cases flagged as suspicious/fraud Tracks customer/business impact
ABR / ACG v1.1 preview Agent bypass patterns; ACG uses envelope compound_confidence on joint approvals Surfaces weak agent/instrumentation gating alongside document BR
TCR / HAR v1.1 preview Tool-call coverage and handoff audit rates on agent-mediated flows Whether logs reconstruct how the evidence package was built

Reference: STANDARD.md (§4.5 preview metrics), METHODOLOGY.md, results_schema.json.

03 Attack levels

Three escalating attack classes

  • L1 — Standard Generation: direct AI-generated documents, no post-processing.
  • L2 — Advanced Diffusion: fine-tuning, editing, metadata manipulation scenarios.
  • L3 — Screen Recapture: synthetic/edited files recaptured through display pipelines.

L3 is a foundation layer in the methodology because recapture can remove or distort provenance cues while preserving plausible visual content — the failure mode most KYC stacks under-report.

04 Audit trails

FRC & agent-era traceability

SDB-26 includes FRC and the FRC A2A Extension (FRC_A2A_EXTENSION.md, v0.5.2) for auditable decisions across human-direct, agent-assisted, and managed-agent channels.

Core links

Highlights in v0.5.2

  • Compound routing combines document FRC with an agent_verdict posture, including INSUFFICIENT × PARTIALLY_ATTESTED → REVIEW and INSUFFICIENT × SUSPICIOUS → ESCALATE — a bad capture plus a risky agent path is not reduced to “upload again”.
  • Normative L0 → agent_verdict mapping so PARTIALLY_ATTESTED is not an informal catch-all.
  • Confidence split: verdict_confidence (core payload) = document layer only; compound_confidence (envelope) = joint compound_verdict; composition IDs (CC_MIN / CC_DOC_ONLY / CC_CUSTOM) support comparable benchmarks.
  • A2A Protocol alignment: optional a2a_correlation and schemas/a2a_v1_surfaces.json follow Task / TaskState shapes from the Agent2Agent specification.
  • Threat model adds T6 (shadow connector / FRC-L0-CONNECTOR-OUT-OF-POLICY) and T7 (opaque secret–workload binding / FRC-L0-SECRET-BINDING-UNKNOWN).

Together this bridges document authenticity to agent-era traceability (instrumentation_trace, L0/L0-D, ABR / ACG / TCR / HAR where applicable).

05 Schemas

Machine-validatable artifacts

  • schemas/frc_schema_v1_1_0.json — current document-layer FRC.
  • schemas/frc_a2a_envelope_v0_2_1.json — audit envelope (agent_verdict, compound_verdict, compound_confidence, optional agent_layer_confidence, a2a_correlation).
  • schemas/a2a_v1_surfaces.json — A2A type surfaces for correlation fields.
  • examples/frc/, scripts/validate_frc_schemas.py — examples and validation.
06 Responsible release

Defender-oriented benchmark

Public artifacts focus on taxonomy, measurement contracts, schema surfaces, and redacted examples that improve defensive evaluation quality.

Operational evasion playbooks and attack-enabling parameter detail are intentionally excluded from open release.

07 Reference implementation

From spec to corpus

  • Forensic packet collection workflow (collect_forensic_packet.py) for repeatable corpus acquisition pipelines.
  • Schema-valid decision artifacts using FRC/FRC A2A outputs and fixtures in this repository.

Related repo artifacts: examples/frc/ · tests/frc/ · CHANGELOG.md

08 Why now

Measurement contract for the agent era

As AI generation quality and agent-mediated onboarding velocity rise, trust controls must move from static checks to measurable, reproducible evidence chains.

SDB-26 provides that measurement contract. Your bypass rate — on your pipeline, not your vendor's marketing slide.