The Provenance Layer Is Not a Truth Machine
A provenance layer is the technical and institutional record that connects an artifact to signed assertions, transformations, embedded signals, displays, and custody. It can show that a claim chain is intact. It cannot establish that the scene, caption, consent, or interpretation is true.
The practical rule is to keep five judgments separate: whether a signal is present, whether it validates, whether its signer is trusted, whether the asserted history is credible, and whether the underlying claim is true.
The New Trust Layer
Every synthetic-media panic eventually reaches for a badge.
The image should say where it came from. The video should reveal whether it was generated, edited, or captured. The model output should carry a machine-readable mark. The platform should label deepfakes. The newsroom should preserve chain of custody. The court should know whether the exhibit has been altered. The voter should not have to become a forensic analyst before deciding whether a clip is real.
This is the world that content provenance and watermarking are trying to build. The Coalition for Content Provenance and Authenticity, or C2PA, defines an open standard for binding signed provenance claims to digital assets. Google’s SynthID takes a different route by embedding an imperceptible signal in outputs made or edited with supported Google AI systems. Visible labels translate selected signals or policy judgments for people. NIST groups provenance data tracking, watermarking, detection, labeling, testing, and auditing within the broader problem of digital content transparency.
Those mechanisms should not be collapsed. A manifest carries assertions. A hard binding uses cryptographic hashes to associate a manifest with an asset. A soft binding, such as a watermark or fingerprint, can help rediscover a manifest after metadata is removed. A trust list helps a validator decide whether to trust a signing credential. A detector reports whether it found a supported signal. A label tells a person what a publisher or platform chose to communicate.
For this essay, the provenance layer means that whole evidence and governance stack: capture records, signed manifests, hashes, watermarks, visible labels, platform displays, detector outputs, source files, edit logs, archive packages, validation software, trust lists, and the policies that decide how each signal is retained, shown, challenged, or ignored. The layers can preserve source and custody. They cannot convert a signed assertion into truth by decorating it with cryptography.
The goal is understandable. Generative AI makes plausible media cheap. It also makes context fragile. A screenshot detaches from a source. A clip loses its upload history. A generated voice crosses from satire into fraud. A campaign image moves through platforms faster than human verification. In that environment, provenance becomes a civic need.
But the central mistake is already visible: treating provenance as if it were truth. It is not. Provenance is a record of claims about origin and handling. Truth is a relationship between a claim and the world.
Current Context
As of August 12, 2026, C2PA 2.4, published in April, remains the current specification. It added the c2pa.ai-disclosure assertion for machine-readable claims about model type, model provenance, scientific domain, and human-oversight level; only modelType is required and the remaining fields are optional. These are assertions supplied by a claim generator, not independent findings that the named model or oversight process was actually used. Version 2.4 also added manifest embedding for HTML and structured text and a JSON-LD-derived view called crJSON. The specification expressly says crJSON is not independently verifiable, an important warning against treating an exported report as the credential it summarizes.
C2PA’s Conformance Program and official Trust List have operated since mid-2025. The program covers generator products, validator products, and certification authorities; its Conformance Explorer exposes the Conforming Products List, C2PA Trust List, and C2PA TSA Trust List. The former Interim Trust List was frozen on January 1, 2026 for legacy support. This is implementation assurance: it helps a validator distinguish a conforming product and accepted signing chain from an unknown one. It does not certify that every assertion is complete or that the depicted event is true.
The legal context has crossed from preparation into application. Article 50 of the EU AI Act has applied since August 2, 2026. The Commission published the final voluntary Code of Practice on June 10, found it adequate on July 8, and published final interpretive guidelines on July 20; the AI Board also assessed the code as adequate. Adherence can support a compliance showing but is not conclusive evidence of compliance. A narrow transition gives systems placed on the market before August 2 until December 2, 2026 to meet the provider-side marking and detection duty in Article 50(2); the other applicable Article 50 duties did not receive that general grace period.
The technical context remains bounded. NIST AI 100-4 treats provenance, watermarking, detection, labeling, prevention, testing, and auditing as complementary approaches rather than a comprehensive solution. Google reported in November 2025 that SynthID had marked more than 20 billion pieces of content and made supported image verification available through Gemini; in December it added video checks across audio and visual tracks. These are provider statements about signals from supported Google systems, not a universal detector and not proof that the depicted claim is false.
The Verification Ladder
A serious provenance review reports which rung of the ladder the evidence reaches.
Present means a manifest, sidecar, soft-bound record, watermark, metadata field, or platform label was found. Presence says nothing yet about cryptographic validity, trust, or factual accuracy.
Valid has a narrower C2PA meaning. The manifest is well formed, remains cryptographically bound to the asset, has an accepted signature and time status, and does not fail the relevant revocation check. A valid manifest can still use a signing credential that the validator does not trust.
Trusted means a valid manifest’s signing credential chains to a trust anchor accepted by that validator. Trust therefore depends on the validator configuration, applicable trust list or private credential store, certificate policy, and time of validation. It does not mean every assertion has been independently corroborated.
Identified is separate again. C2PA 2.x primarily represents the machine identity of the application, service, or hardware product signing a claim. Human or organizational identity requires additional assertions, such as the Creator Assertions Working Group profiles recommended by C2PA. A product’s trusted signing certificate should not silently become proof that a named photographer, publisher, or eyewitness made the content.
Credible means the asserted history survives comparison with the native file, ingredient manifests, source testimony, edit records, tool behavior, and institutional process. C2PA also distinguishes created_assertions, which are attributed to the signer, from gathered_assertions, which the claim includes but does not attribute to that signer. A viewer that flattens those categories overstates who vouched for what.
True is the final and different question: does the artifact or caption accurately represent the world? A valid and trusted credential can accompany a staged scene, misleading crop, false caption, nonconsensual recording, or incomplete account. An uncredentialed file may be authentic because its toolchain never supported the standard or stripped the manifest.
This is real value. Synthetic media does not only deceive by being false. It deceives by being untethered. The viewer sees a surface without knowing whether it is capture, reconstruction, satire, propaganda, simulation, evidence, or decoration. Provenance gives the artifact a path.
But the path is not the destination. A valid credential can accompany a misleading caption. A camera can faithfully record a staged event. A credentialed publishing workflow can vouch for a photo that frames reality selectively. A model can generate a labeled synthetic scene that is ethically legitimate. A file can lack credentials because an old tool stripped metadata, not because it is fake. A bad actor can circulate false material with a technically valid provenance chain.
C2PA's own explainer is careful about this distinction: provenance can help establish origin and history, but it cannot by itself determine whether digital content is true, accurate, or factual. That sentence should be treated as the constitutional warning for the whole field.
Watermarking and Labels
Watermarking solves a different part of the problem. If embedded metadata can be removed by uploads, screenshots, recompression, and format conversion, a signal carried in the content may survive longer. Google says SynthID covers generated text, images, audio, and video across supported systems. Its Gemini verification surfaces check for that provider-specific signal and can identify likely marked regions or time segments. The result should say SynthID detected, not collapse that observation into fake.
That matters because the provenance layer has to survive hostile and careless environments. Most people do not preserve original files. Most platforms transform media. Most users encounter copies of copies. A system that only works in ideal chain-of-custody conditions will help newsrooms and archives more than public feeds.
Still, watermarking has its own boundary. A detector can usually say something narrower than the public wants. It may say that a particular watermark was detected, not that the media is fake. It may say that no known watermark was found, not that the media is human-made. It may work for outputs from one ecosystem, not all generators. It may degrade under editing, transcription, paraphrase, model laundering, analog capture, or adversarial removal.
That weakness has experimental support, but the scope matters. In the 2024 revision of Robustness of AI-Image Detectors, Mehrdad Saberi and colleagues found an evasion-versus-spoofing tradeoff for low-perturbation image watermarks under diffusion purification, demonstrated a model-substitution attack against higher-perturbation methods, and constructed black-box spoofing attacks that made unwatermarked images test as marked. The paper does not establish that every later proprietary, multimodal, or high-distortion scheme fails under every condition. It establishes the governance burden: name the watermark version, media type, attack model, transformation suite, detector threshold, and observed false-positive and false-negative rates before calling it robust.
Visible labels add another layer. A platform can mark an image as AI-generated, manipulated, disputed, or lacking context. A creator can disclose synthetic elements. A publisher can explain the provenance chain in ordinary language. The Partnership on AI's synthetic-media framework treats both direct disclosure, such as labels and disclaimers, and indirect disclosure, such as watermarking and cryptographic provenance, as useful practices.
A product interface can introduce another ambiguity by combining raw detection with model-generated explanation. Preserve the detector’s actual result, supported formats, tool version, threshold or confidence representation, and any abstention separately from the prose shown to the user. Otherwise the explanatory model can become an undocumented second detector.
The right model is layered evidence. Credentials, watermarks, visible labels, institutional reputation, original files, editorial process, forensic analysis, source interviews, and public correction channels should reinforce one another. None should pretend to be the whole truth system.
The Article 50 Problem
Law is now turning synthetic-media transparency into an operational requirement.
Article 50(2) requires providers of systems that generate synthetic audio, image, video, or text to mark outputs in a machine-readable format and make them detectable as artificially generated or manipulated. The technical solution must be effective, interoperable, robust, and reliable as far as technically feasible, taking account of media limits, implementation cost, and the generally acknowledged state of the art. The duty does not apply to the extent a system performs an assistive function for standard editing or does not substantially alter the deployer’s input or its semantics.
Article 50(4) places a different duty on professional deployers: clearly disclose deepfake image, audio, or video and certain AI-generated or manipulated text published to inform the public on matters of public interest. Public-interest text is exempt where it has undergone substantive human review or editorial control and a person holds editorial responsibility. Evidently artistic, creative, satirical, fictional, or analogous works receive an adapted disclosure duty that should not hamper their display or enjoyment. The Commission guidelines say a deployer cannot satisfy the human-facing deepfake duty merely by relying on the provider’s hidden machine-readable mark; disclosure must be clear and perceivable by first exposure. Article 50(5) also requires the information to be clear, distinguishable, and accessible.
This is a major shift. The question is no longer only whether a company thinks labels are good practice. The question becomes how an entire media environment implements marking, detection, and disclosure at scale while keeping the provider’s technical duty separate from the deployer’s communication duty. The voluntary code gives signatories an approved route for demonstrating compliance, while non-signatories may use other adequate means.
That sounds simple until the artifact becomes mixed. A journalist uses AI transcription to summarize a real interview. A designer uses generative fill on a human photograph. A campaign edits background noise out of a real speech. A student uses a model to translate a human-written essay. A publisher uses an AI tool to resize, crop, sharpen, and caption a real image. A documentary reconstructs a historical scene and labels it in the credits. An artist intentionally blurs the line between capture and synthesis.
Binary labels collapse these cases. “AI-generated” can mean wholly synthetic, materially transformed, translated, expanded, filled, summarized, or lightly assisted. The legal exceptions make the distinction consequential, not merely semantic. If the label is too broad, it becomes noise. If it is too narrow, it misses meaningful manipulation. If it is too technical, ordinary users cannot act on it. If it is moralized, legitimate art, accessibility work, and ordinary editing get treated as contamination.
This is the governance problem hidden inside Article 50-style transparency. Disclosure is necessary, but it must describe the relevant transformation, identify who owes which duty, survive the publication interface, and remain accessible. Otherwise the mark becomes compliance theater: machine-readable enough for a pipeline, ambiguous enough for people, and weak enough for a platform to convert into decorative certainty.
Failure Modes
The first failure mode is false absence. Users may learn to treat unlabeled content as real. That is the wrong inference. Absence of a credential, watermark, or label may mean no participating tool was used, the signal was stripped, the platform failed to display it, the detector missed it, or the content came from an ecosystem outside the trust layer.
The second is false presence. A technically valid provenance record can accompany a misleading artifact. The signature attributes created assertions to a signer; it does not prove that the signer is honest, authorized to depict the subject, contextually complete, or correct. A spoofed watermark can create a still narrower false signal.
The third is trust monopoly. If the public learns to trust only credentials from large platforms, device makers, model labs, or publishers, provenance infrastructure can centralize authority over what counts as authentic. C2PA’s own harms analysis notes participation barriers for civic and independent media and warns that uncredentialed authentic material should not be dismissed. Journalists, activists, artists, and human-rights witnesses may lack access to conforming tools or may need anonymity.
The fourth is privacy leakage. A manifest or repository lookup can reveal workflow, time, device relationship, editing tool, location, publication chain, or identity. The C2PA harms model warns that sensitive data can be added inadvertently and that redacted information may remain discoverable through earlier remotely stored manifests. It also identifies surveillance and human-rights risks in hostile legal or political settings. Credential creation, remote lookup, and viewer telemetry therefore need separate data-minimization review.
The fifth is validation drift. A platform may preserve the asset but not the validator version, trust-list snapshot, certificate status, time-stamp result, or raw status codes used for a past decision. A later reviewer then sees the same file through different trust infrastructure and cannot reconstruct why an earlier badge appeared.
The sixth is ritual reassurance. A badge appears. The viewer relaxes. The platform has performed trust. The institution has performed diligence. But the underlying question may remain unanswered: what does this media claim, and what evidence supports that claim?
The seventh is recursive pollution. Synthetic content that lacks durable provenance can re-enter search indexes, training datasets, retrieval systems, and public archives. Future systems then summarize, cite, remix, and normalize media whose origin has dissolved. Provenance is not only about today’s viewer. It is about the memory available to tomorrow’s machines.
The Governance Standard
A serious provenance regime should start with precise claims.
First, name the signal and status. “Captured by camera,” “AI-generated,” “AI-edited,” “valid manifest,” “trusted signer,” “watermark detected,” “publisher-verified,” and “unknown origin” are different claims. Compressing them into one badge makes the interface cleaner and the institution less accurate.
Second, preserve negative-result uncertainty. Platforms should not imply that undetected means authentic. The public-facing language should say what was actually checked: no supported credential found, no known watermark detected, manifest invalid, signer not trusted by this validator, provenance unavailable, or source not independently verified.
Third, retain a validation receipt. For every consequential check, preserve the asset hash, manifest or sidecar, retrieval path, validator and version, trust-list source and snapshot, certificate and time-stamp status, raw validation codes, watermark detector and version, transformation applied before testing, human-facing label, and the time of the decision. The receipt should distinguish an embedded credential from a remotely recovered one and a machine result from an explanatory summary.
Fourth, preserve chain of custody. Courts, newsrooms, human-rights investigators, election officials, and public agencies should retain native files, ingredients, edit and export logs, publication records, platform displays, and verification notes. A badge on a compressed social-media copy is not enough. This is the same record discipline required when synthetic evidence enters a court record.
Fifth, make provenance contestable. Users and institutions need notice, correction, and appeal paths when content is mislabeled, a credential is stripped, a trust decision changes, a platform misreads a signal, or a label damages legitimate speech. Preserve the challenged version and decision record before changing the display.
Sixth, minimize creator and viewer data. Do not put identity, location, device, or workflow details into public assertions merely because the schema permits them. Explain remote-manifest lookups and viewer telemetry, offer a safe local or anonymous verification path where feasible, and ensure that redaction reaches repositories and cached copies.
Seventh, test the whole publication path. Procurement should cover export, resizing, screenshots, recompression, transcoding, CDN processing, social upload, archiving, accessibility conversion, manifest recovery, revocation, and offline viewing. Record which transformations preserve, invalidate, or remove each signal. A credential that disappears at the first handoff is an archive feature, not a public label.
Eighth, protect plurality. C2PA, watermarks, platform labels, editorial notes, and legal duties should interoperate where possible, but no vendor, trust list, or state should become an oracle of authenticity. Institutions should accept alternative chains of custody and must not downgrade witnesses merely because they lack conforming hardware or safe access to identity credentials.
Ninth, treat detectors as leads. A watermark detector, synthetic-content classifier, or credential viewer should trigger proportionate review rather than an automatic verdict. Validate it on the relevant medium and transformations, publish known error and abstention behavior, and forbid unsupported inferences from a negative result. This links the provenance problem to detector-driven discipline.
Tenth, review incidents and correct the record. Mislabeling, stripped credentials, spoofed watermarks, broken verification displays, trust-list failures, privacy leakage, and inaccessible disclosures should enter an incident-reporting process with owners, containment, affected-party notice, correction, recurrence testing, and public explanation where warranted.
What This Changes
The synthetic age does not only create false things. It creates orphaned things.
A voice arrives without a throat. A photo arrives without a camera. A citation arrives without a source trail. A clip arrives without the minutes before and after it. A generated image enters the feed as if it had fallen from the sky. The artifact becomes pure surface, and the surface asks to be believed.
Provenance is the effort to attach memory back to the artifact. It says: this came from somewhere, passed through something, was changed by someone or some system, and carries a record that can be examined. That is an institutional good.
But provenance can also become another interface of obedience. The user sees a badge and stops asking questions. The platform shows a label and claims responsibility has been discharged. The state mandates a mark and calls the epistemic crisis managed. The credential becomes a symbol of control standing in for control itself.
The better discipline is colder and more useful. Follow the credential. Read the claim. Ask what is missing. Preserve originals. Label uncertainty. Protect private creators. Separate origin from truth. Treat synthetic media governance as evidence infrastructure, not a priesthood of badges.
A provenance layer is useful because copied, transformed, and synthetic artifacts routinely lose source context. It becomes dangerous when it pretends to end interpretation. The badge should begin the investigation, not replace it.
Source Discipline
Provenance sources have different authority. The normative C2PA specification defines formats, validation, and trust semantics. Its explainer, identity recommendation, and harms model are informative accounts of intended scope and known risks. The Conformance Program supports product and certificate assurance, not factual truth. NIST AI 100-4 is technical-governance guidance, not law. Article 50 and its 2026 amendment are legal text; Commission guidelines and FAQs explain the regulator’s interpretation. The code is voluntary and has been assessed as adequate, but the Commission expressly says adherence is not conclusive evidence of compliance.
Google’s SynthID posts are primary evidence for Google’s own deployment claims, supported media, and interface behavior; they do not validate third-party generators or establish field-wide error rates. The Saberi et al. paper is primary research on specified image detectors and attacks, not a finding about every watermark released later. The Partnership on AI framework is a voluntary practice framework, not a standard or legal obligation.
Claims should state the evidence level precisely. “Credential present” means a supported manifest was located. “Manifest valid” means a particular validator completed defined structural and cryptographic checks. “Signer trusted” means the credential was accepted under a named trust configuration. “Watermark detected” means a named detector found its supported signal under recorded conditions. “AI-generated” requires a definition of generation or material manipulation. “No signal found” means only that the check found no supported signal, not that the artifact is human-made, authentic, consensual, or true.
For high-stakes media, cite and preserve the artifact pathway: native file, hash, C2PA manifest or sidecar, ingredient chain, verification result and raw status, validator and trust-list version, watermark result, transformation history, platform label, source publication, upload context, reporter notes, legal or editorial use, and correction history. Treat a social screenshot, platform label, detector score, signed assertion, and legal disclosure as different artifacts with different evidentiary value under the site’s research-integrity standard.
Related Pages
- Content Provenance and Watermarking defines the wider technical field and its boundary tests.
- Provenance and Content Credentials turns the distinction into a publishing and archive protocol.
- The Synthetic Evidence Becomes the Court Record applies chain-of-custody discipline to legal evidence.
- The AI Detector Becomes the Discipline Machine examines the danger of converting a fallible classifier into a verdict.
- The Authenticity Debt Becomes the Trust Ledger treats provenance, detection, disclosure, and human verification as separate trust layers.
- The Takedown Button Becomes Synthetic Media Governance follows synthetic-media governance from disclosure into remedy and appeal.
- AI Data Provenance extends the record from published media into datasets, models, and downstream use.
- AI Incident Reporting supplies the correction and recurrence process when provenance controls fail.
Sources
- C2PA, Content Credentials: C2PA Technical Specification 2.4, April 2026, for terminology, version history, bindings, AI Disclosure, validation, and trust semantics.
- C2PA, C2PA and Content Credentials Explainer, version 2.4, for provenance limits, removal, identity scope, and the distinction between created and gathered assertions.
- C2PA, Human and Organizational Identity Recommendation, version 2.4, for the separation of machine signer identity from human and organizational identity assertions.
- C2PA, Harms Modelling, version 2.4, for accessibility, privacy, surveillance, independent-media, and uncredentialed-content risks.
- C2PA, Conformance Program, reviewed August 12, 2026, for conforming-product and trust-list governance.
- NIST, Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency, November 20, 2024, updated April 8, 2026.
- European Union, Regulation (EU) 2024/1689, Artificial Intelligence Act, Article 50, and Regulation (EU) 2026/1744, for the limited transition applicable to pre-August 2 systems.
- European Commission, Guidelines on Transparency Obligations for Providers and Deployers of Certain AI Systems, published July 20, 2026.
- European Commission, Code of Practice on Transparency of AI-Generated Content, final code published June 10, 2026.
- European Commission, Opinion on the Assessment of the Code of Practice, published July 9, 2026.
- European Commission, Transparency Obligations Under Article 50 of the AI Act: Questions and Answers, reviewed August 12, 2026, for scope, disclosure, exceptions, compliance routes, and the transition period.
- Partnership on AI, Responsible Practices for Synthetic Media: A Framework for Collective Action, February 27, 2023.
- Google, How We’re Bringing AI Image Verification to the Gemini App, November 20, 2025, and You Can Now Verify Google AI-Generated Videos in the Gemini App, December 18, 2025.
- Mehrdad Saberi et al., Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks, arXiv:2310.00076v2 [cs.CV], revised February 14, 2024.