A Conservation-of-Trace Framework for Authenticating Imagery in the Age of Generative AI
Guest Blogger: Khaled S. Al Sannat
Generative models have dissolved the oldest working assumption of visual evidence: that a photograph is, by default, a witness. The reflex of the field has been to ask a binary question — is this image real or fake? — and to chase it with classifiers that are accurate today and obsolete with the next model release. This paper argues that the binary question is the wrong one, and proposes a different foundation. An authentic capture is the end of a physical chain — photons, optics, a color filter array, a silicon sensor, an analog-to-digital converter, an interpolation pipeline, a compressor — and every link in that chain imprints a statistical trace that cannot be wished away. A synthetic artifact has no such chain; it can only approximate the trace, at a cost, and never recover what laundering has destroyed. We formalize this as the Conservation of Trace and translate it into the language a courtroom understands: detection as hypothesis testing with a known error rate, attribution as estimation with a hard precision floor, and laundering as a channel that can only erode provenance, never fabricate it. From these bounds we derive an operational, layered methodology and a single discipline that separates serious work from theater: never report a verdict, report a calibrated likelihood with its uncertainty.
1. The Collapse of the Image as Self-Evident Witness
For a century the photograph carried an implicit warrant. It was treated as a trace of something that had been in front of a lens, and the burden of proof sat with anyone claiming otherwise. That warrant is gone. A practitioner can now generate, in seconds and for nothing, a face that never lived, a document that was never signed, a scene that never happened — at a fidelity that defeats the unaided eye and, increasingly, the trained one.
The instinctive response has been to build a better detector: train a network on examples of the latest generator and let it flag the fakes. This is a losing posture, and it is worth being precise about why. A detector trained on the outputs of one family of models learns the fingerprints of that family. When a new architecture appears — and one appears every few months — the learned fingerprints no longer apply, and accuracy collapses on exactly the material the analyst most needs to judge [3]. The detector is always fighting the last war.
The deeper error is not technical but conceptual. “Is it fake?” demands a verdict from an artifact that, on its own, cannot supply one. The question that an artifact can answer is different: what is the budget of provenance information it carries, and with what certainty can that budget be read? Reframing the problem this way moves the discipline from divination toward accounting and accounting is something a court, an opposing expert, and an appellate judge can all interrogate.
The organizing principle of this paper is what we have elsewhere called the Conservation of Trace: computation and representation leave an inescapable physical footprint, bounded by physical law, that can be destroyed but not conjured. The earlier statement of that principle lived in the storage domain — why a securely deleted file is never quite gone at the level of the medium. The same principle, carried into the imaging domain, is the foundation of authentication. An honest capture cannot help leaving its trace; a forgery cannot help lacking the authentic one, and cannot manufacture it from nothing.
The contribution of this paper is threefold. First, we reframe image authentication from a binary verdict into an exercise in provenance-information accounting — bounded above by the distinguishability of the competing hypotheses and below by the precision of attribution. Second, we express the Conservation of Trace through four established results — the likelihood-ratio test, Stein’s lemma, the Cramér–Rao bound, and the data-processing inequality — and show that together they render laundering a one-way operation: provenance can be eroded but never fabricated. Third, we translate these bounds into a layered, admissible methodology and a single reporting discipline — report a calibrated likelihood with its uncertainty, never a verdict.
2. The Physical Provenance Chain
To understand what a forgery is missing, one must first inventory what an authentic image cannot avoid carrying. A real photograph is the residue of a deterministic physical pipeline, and each stage deposits a signature.
2.1 The sensor fingerprint (PRNU)
No two image sensors are identical. Microscopic, unavoidable variation in the manufacture of silicon photosites means each pixel responds to light with a slightly different gain. The dominant and most stable component of this is photo-response non-uniformity (PRNU) — a multiplicative noise pattern that is effectively a hardware fingerprint, unique to an individual sensor and stable across its lifetime [1]. PRNU has been the backbone of source-camera identification for nearly two decades, and the detection of a reference pattern in a questioned image is properly posed not as a heuristic but as a statistical decision problem [2]. The forensic significance for synthetic media is blunt: a fully generated image was never exposed to any photosite array, so there is no authentic PRNU tied to any claimed device. A face swapped onto real footage locally suppresses or disrupts the sensor pattern in precisely the manipulated region.
2.2 The color filter array and demosaicing
Most sensors are monochrome detectors overlaid with a periodic color filter — the Bayer array — so that each photosite natively records only one of red, green, or blue. The missing two channels at every pixel are interpolated by a demosaicing algorithm. This interpolation is not free of consequence: it imposes specific, periodic correlations between neighboring pixels and across color channels. Authentic camera output carries these demosaicing correlations as a matter of physics; a generative pipeline that never passed through a real CFA produces output whose inter-channel structure is absent, inconsistent, or characteristic of the generator rather than of any camera.
2.3 Compression genealogy
Real-world images almost always have a compression history, and lossy compression leaves traces. The quantization of frequency coefficients imprints statistical regularities, and recompression — the ordinary fate of any image that has been edited and re-saved — leaves a detectable double-compression signature. The compression genealogy of an artifact is therefore evidence in itself: a file whose claimed history is “straight off the camera” but whose coefficient statistics betray a second compression has already told the analyst something. Synthetic images have a different, and frequently absent, compression lineage. None of these signatures is, in isolation, decisive. Their force is cumulative and physical: together they constitute the positive evidence of an authentic capture, the thing a forgery must counterfeit in full and in mutual consistency.
3. The Signatures of Synthesis
Where the previous section described what authenticity leaves behind, this section describes what synthesis leaves behind — and what it conspicuously omits. The two are complementary halves of the same accounting.
3.1 Spectral artifacts of generation
Generative networks build images by repeatedly up-sampling from a low-dimensional representation to full resolution. These up-sampling operations are not spectrally neutral; they deposit periodic artifacts that are largely invisible in the spatial domain but glaring in the
frequency domain. A frequency-domain analysis of GAN output reveals severe, regular artifacts that are consistent across architectures, datasets, and resolutions, and that trace directly to the up-convolution operations common to the entire model class [4]. This is a structural property of how such images are made, not an incidental flaw of one model. Diffusion models leave their own distinct residual signatures, which a forensic analysis can likewise surface [5].
A crucial and honest caveat travels with these findings. The same body of work that showed CNN-generated images are “surprisingly easy to spot” qualified the claim with two words: for now [3]. Spectral and statistical signatures are real, but they are model-dependent and can be
attenuated by an adversary who is aware of them. They are evidence to be weighed, not oracles.
3.2 Violations of physics and physiology
The most durable signatures are the ones rooted in physical law rather than in the quirks of a particular network, because they constrain any image that purports to depict a real scene. Three families are well established:
- Illumination consistency. In a genuine photograph the shading and shadows across all objects are consistent with one lighting environment. Geometric techniques can recover constraints on the light-source position from cast and attached shadows and from surface shading, and an arrangement that admits no consistent solution is physical evidence of compositing or fabrication [7, 8].
- Reflection geometry. Reflections — in mirrors, in water, in the cornea of an eye — obey perspective. Inconsistent reflections that cannot be reconciled with a single coherent geometry are a tell that the elements of a scene were never co-present [9].
- Physiological signal. A real human face, recorded on video, carries a faint pulse: blood flow modulates skin color in a way that remote photoplethysmography can recover. Synthetic faces frequently fail to reproduce a spatially and temporally coherent pulse, and
this biological signal has been used as an implicit descriptor of authenticity that pure
pattern-matching neglects [6].
These constraints share a property that makes them valuable: they bind the forger to the laws of optics and biology, not merely to the statistics of a training set. They are harder to evade because evading them means satisfying physics, not fooling a classifier.
4. The Mathematical Core: From Verdict to Bounded Attribution
Everything above becomes a methodology only when it is expressed in the language of measurement and uncertainty. Four results carry the weight. They are not novel mathematics; their value is that they convert a forensic intuition into statements that survive crossexamination.
4.1 Detection is hypothesis testing
Authentication is a decision between two hypotheses: H₀, that the artifact is an authentic capture from the claimed pipeline, against H₁, that it is synthetic. The optimal decision at a fixed false-alarm rate is, by the Neyman–Pearson lemma, the likelihood-ratio test [12]:
Λ(x) = p(x | H1) / p(x | H0) declare synthetic when Λ(x) ≥ τ
The threshold τ is not a matter of taste; it fixes the trade-off between false accusations and missed forgeries, and it must be stated. This single reframing already disciplines practice: it forces an explicit error budget where a classifier output offers only a label.
4.2 There is a ceiling on detectability
How well any detector can possibly perform is bounded by how distinguishable the two hypotheses are. By Stein’s lemma, for a fixed false-alarm rate the best achievable exponent of the miss probability is the Kullback–Leibler divergence between the authentic and synthetic
feature distributions [13]:
best error exponent = DKL( P0 ‖ P1 )
The consequence is sobering and clarifying. As generators improve, the synthetic feature distribution converges toward the authentic one, the divergence shrinks toward zero, and the achievable separation collapses — for every detector, not merely for the ones we have today.
This is why a pure-detection strategy is structurally doomed at the limit, and why provenance and physical constraint, which do not depend on a residual statistical gap, must carry the load.
4.3 There is a floor on attribution precision
When the analyst moves from “is it synthetic?” to “did it come from device D?”, the relevant quantity is the precision of an estimate — for instance, the strength of a PRNU correlation. The Cramér–Rao bound states that the variance of any unbiased estimator of a parameter θ is at least the inverse of the Fisher information [11]:
Var(θ̂) ≥ 1 / IF(θ)
This is the courtroom’s most lethal question rendered as mathematics: with what precision can this attribution be asserted? The bound gives an honest ceiling on confidence dictated by the available signal, the noise, and the sample. An expert who can quote it is doing measurement; one who cannot is offering an opinion.
4.4 Laundering can only destroy provenance, never create it
Re-compression, resizing, filtering, and re-capture are the ordinary laundering operations that erode evidence. Their effect is governed by the data-processing inequality. Modeling origin as X, the captured artifact as Y, and the laundered artifact as Z in the chain X → Y → Z, no
processing step can increase the information the artifact carries about its origin [13]:
I( X ; Y ) ≥ I( X ; Z )
This is the Conservation of Trace stated exactly, and it cuts in both directions. It bounds what a defender can hope to recover from a heavily processed file — honesty about that limit is itself part of the method. But it equally forbids a forger from manufacturing genuine provenance by processing: one cannot add an authentic sensor fingerprint by passing an image through a filter. Trace is conserved in the only sense that matters forensically — it leaks away, but it does not spring from nothing.
4.5 The forgery-cost asymmetry
The strategic payoff of this framework follows from combining the layers. The sensor, the CFA, the compression history, the generation spectrum, the illumination, and the physiology are approximately independent constraints, each arising from a different physical or statistical mechanism that no generator jointly optimized. To pass all of them simultaneously and consistently, an adversary must satisfy a set of coupled constraints at once. No single layer is unbreakable, and a sophisticated adversary can target several; but because the layers are physically distinct, defeating them does not get cheaper in aggregate — it compounds. The independence is approximate rather than exact — a single physical re-capture can launder several layers at once — and Section 6 examines that coupling directly. Authentication built as defense-in-depth across independent trace layers raises the cost of a clean forgery faster than it raises the cost of analysis. That asymmetry, not any one detector, is the durable advantage.
4.6 Calibration: from scores to likelihoods
The four bounds describe what is achievable; they do not, by themselves, tell an analyst how to produce a likelihood ratio for a particular exhibit. In practice no layer yields a closed-form likelihood p(x | H). Each layer produces a score, and the score is mapped to a calibrated
likelihood ratio by empirical calibration on labelled validation data — the same score-tolikelihood-ratio calibration, typically logistic, that forensic science already uses under the evaluative-reporting framework [16, 17]. The tractable statistical layers carry their own nearclosed-
form null distributions and need little calibration: the PRNU decision, for instance, is read from the peak-to-correlation-energy statistic, whose behaviour under the no-match hypothesis is known [2]. The geometric and physiological layers are where intuition rebels — one does not solve the optics of a reflection in a cornea in closed form. What the analyst measures there is a geometric or temporal inconsistency residual, and what is reported is that residual’s calibrated weight against a reference population, not an analytic probability. This is the join between the theory and the field: a constraint that cannot be computed exactly can still be measured, calibrated, and given a defensible weight — which is what keeps the method operational rather than ornamental. A fair objection is where the labelled data for the physical layers comes from, since no archive of authentic-versus-forged corneal reflections sits ready on a shelf. Two sources answer it. For the geometric constraints, physically-based rendering supplies ground truth directly: a scene composited under a known, consistent light field furnishes the authentic class, and the same scene perturbed into a deliberately impossible configuration furnishes the forged one. Where simulation is impractical — a coherent facial pulse, for instance — the forged class can be drawn from generative models themselves, an increasingly standard expedient whose one caveat is the familiar one: data minted from known generators calibrates against
known generators, which is exactly why the physical layers are never asked to carry a case by themselves.
5. An Operational Methodology for Investigation Units
The bounds above prescribe a workflow. It is layered, it is explicit about uncertainty, and it is built to be admissible rather than merely persuasive.
5.1 The layered analytical stack
Each layer answers a narrow question and contributes evidence to a final, fused statement. No layer is treated as conclusive on its own.
The first layer deserves emphasis because it is the only one that can deliver positive proof rather than inference. Cryptographic content provenance — the C2PA Content Credentials standard — binds a signed manifest of origin and edit history to a file, verifiable offline with no
central database, and tamper-evident in the sense that any alteration breaks the signature [14]. What a valid credential proves, however, is narrower than it first appears: that a named signer vouched for this provenance and that the bits have not changed since signing — not that the depicted scene is real. A credential is therefore only as strong as its signer and the capture pipeline behind it, and its weight must scale with the root of trust: a hardware-attested capture credential is strong evidence, whereas a software credential generated on an open device —where manipulated pixels can be fed to a legitimate signer before signing — is cryptographically valid yet forensically hollow. The discipline lies in the converse: the absence of a credential proves nothing, because the overwhelming majority of legitimate media still carries none, and stripping a credential is trivial. Provenance thus speaks in three registers, not two — strong when hardware-anchored and present, weak when merely software-signed, and silent when absent — and an investigator who collapses them will mislead a court. Reading that root of trust is itself a specialized skill — telling a hardware-attested credential from a software-signed one means validating a certificate chain against a trust list, which most police investigation units are not staffed to do. The fail-safe default, absent that cryptographic competence, is to treat any credential whose root cannot be verified as a weak L0 rather than a strong one.
5.2 The reporting doctrine
This is the single practice that distinguishes court-grade work from a confident guess. The output of the analysis is never the words “real” or “fake.” The output is a calibrated likelihood ratio — how much more probable the evidence is under the synthetic hypothesis than under the
authentic one — accompanied by a stated error rate and an explicit note on how far laundering may have degraded the available signal. A binary verdict hides its uncertainty; a likelihood statement exposes it, which is exactly what makes it defensible. This doctrine is not merely good taste; it is what admissibility requires. The Daubert standard for expert testimony asks, among other things, for the known or potential error rate of a technique and whether it rests on testable, peer-reviewed method [15]. A pipeline that reports likelihood ratios with quantified error rates answers that demand directly. A classifier that emits a label cannot. The rigor is not a constraint bolted onto the method — it is the method’s reason for being trusted.
5.3 Triage: managing the cost of analysis
The stack is not run exhaustively on every artifact. Estimating PRNU, analysing DCT statistics, fitting three-dimensional illumination models, and recovering an rPPG pulse for thousands of exhibits in a single case is rarely affordable, so the layers are deployed as a cost-ordered cascade. Cheap, high-throughput screens come first and run at population scale: the presence of a C2PA credential, a coarse frequency-domain check, and a metadata-and-compressionhistory pass cost almost nothing per file. Exhibits that those screens flag — or that the case makes material — are escalated to the expensive layers, device-specific PRNU correlation and physical-scene modelling, which run on a prioritized subset rather than the whole population.
The cost asymmetry of Section 4.5 concerns the forger’s budget; this cascade is how the analyst keeps their own budget bounded while preserving the option to bring every layer to bear on the exhibits that decide a case.
5.4 Extension to moving imagery
Most contested material today is video — defamation, extortion, and fabricated news rarely arrive as a single still — and video does not weaken this framework; it strengthens it. Every layer above acquires a temporal dimension, and temporal redundancy adds constraints rather than removing them: the sensor fingerprint must stay consistent frame to frame, the compression history now includes a group-of-pictures structure and inter-frame prediction whose re-encoding leaves its own traces, and the physiological pulse must be coherent in time as well as space. A new layer also appears — audio-visual synchronization, where a manipulated face and its soundtrack drift out of phoneme-to-viseme alignment. Each added constraint is one more independent thing the forger must satisfy at once, so the forgery-cost
asymmetry compounds further for video, at a correspondingly higher analysis cost that the triage cascade above is meant to absorb. The accounting is identical; there is simply more trace to account for.
6. The Limits of the Method — and Why They Are Its Strength
A position paper that claimed a solution would be the least credible kind. The honest boundaries of this approach are precisely what make it scientific, and they should be stated to any tribunal before an opponent does.
First, learned signatures generalize poorly. Detectors tuned to known generators degrade on unseen ones, and even careful analyses of the state of the art conclude that detection is far from settled [3, 10]. This is why the framework leans on physical and provenance layers that do
not depend on a particular model.
Second, the analog hole is real. An adversary who displays a synthetic image and photographs it with a genuine camera can wrap a forgery in an authentic sensor fingerprint and a legitimate compression history. The data-processing inequality explains why this is powerful — and also bounds it, because the re-capture cannot reproduce the physical and physiological consistency of an actual scene, which the L4 layer is built to test.
Third, every passive signal can be attacked by an adversary who understands it, and even cryptographic provenance survives only where it was applied at capture and not stripped downstream. The sharpest instance is the fingerprint-copy attack, in which an adversary
estimates a target camera’s PRNU from its images and superimposes it to frame an innocent owner; even this leaves its own residue and is detectable, and planting a sensor fingerprint without a trace is far harder than it first appears [18]. There is no permanent victory here. There is only a discipline that reports what it knows, quantifies what it does not, and raises the cost of deception faster than the cost of detection.
These limits are the argument, not an apology for it. A method that names its own error rate is one a court can rely on; a method that promises certainty is one it should not.
Conclusion: Accounting, Not Divination
The generative surge has not made visual evidence worthless; it has made the casual reading of it untenable. The path forward is not a better lie-detector but a change of question. An authentic image is the residue of a physical chain that leaves an inescapable trace; a synthetic one lacks that trace and cannot fabricate it; and laundering can erode the trace but never invent it. Read this way, authentication becomes an exercise in provenance-information accounting, bounded above by the distinguishability of the hypotheses, bounded below by the precision of attribution, and governed throughout by the conservation of the trace itself. For an investigation unit, the practical inheritance is concrete: build the analysis in independent physical layers, anchor it in cryptographic provenance where it exists, and report a calibrated likelihood with its uncertainty rather than a verdict. The discipline will not make every case decidable. It will make every conclusion defensible — which, in the place where this work ultimately lands, is the only standard that counts.
References
[1] Lukáš, J., Fridrich, J., & Goljan, M. (2006). Digital Camera Identification from Sensor Pattern
Noise. IEEE Transactions on Information Forensics and Security, 1(2), 205–214.
[2] Chen, M., Fridrich, J., Goljan, M., & Lukáš, J. (2008). Determining Image Origin and Integrity
Using Sensor Noise. IEEE Transactions on Information Forensics and Security, 3(1), 74–90.
[3] Wang, S.-Y., Wang, O., Zhang, R., Owens, A., & Efros, A. A. (2020). CNN-Generated Images
Are Surprisingly Easy to Spot… for Now. Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition (CVPR), 8695–8704.
[4] Frank, J., Eisenhofer, T., Schönherr, L., Fischer, A., Kolossa, D., & Holz, T. (2020). Leveraging
Frequency Analysis for Deep Fake Image Recognition. Proceedings of the 37th International
Conference on Machine Learning (ICML), PMLR 119, 3247–3258.
[5] Corvi, R., Cozzolino, D., Zingarini, G., Poggi, G., Nagano, K., & Verdoliva, L. (2023). On the
Detection of Synthetic Images Generated by Diffusion Models. ICASSP 2023 — IEEE
International Conference on Acoustics, Speech and Signal Processing, 1–5.
[6] Ciftci, U. A., Demir, İ., & Yin, L. (2020). FakeCatcher: Detection of Synthetic Portrait Videos
using Biological Signals. IEEE Transactions on Pattern Analysis and Machine Intelligence.
doi:10.1109/TPAMI.2020.3009287.
[7] Kee, E., O’Brien, J. F., & Farid, H. (2013). Exposing Photo Manipulation with Inconsistent
Shadows. ACM Transactions on Graphics, 32(3), 28:1–28:12.
[8] Kee, E., O’Brien, J. F., & Farid, H. (2014). Exposing Photo Manipulation from Shading and
Shadows. ACM Transactions on Graphics, 33(5), 165:1–165:21.
[9] O’Brien, J. F., & Farid, H. (2012). Exposing Photo Manipulation with Inconsistent Reflections.
ACM Transactions on Graphics, 31(1), 4:1–4:11.
[10] Gragnaniello, D., Cozzolino, D., Marra, F., Poggi, G., & Verdoliva, L. (2021). Are GAN
Generated Images Easy to Detect? A Critical Analysis of the State-of-the-Art. IEEE
International Conference on Multimedia and Expo (ICME), 1–6.
[11] Kay, S. M. (1993). Fundamentals of Statistical Signal Processing, Volume I: Estimation Theory.
Prentice Hall. (Cramér–Rao lower bound; Fisher information.)
[12] Kay, S. M. (1998). Fundamentals of Statistical Signal Processing, Volume II: Detection Theory.
Prentice Hall. (Neyman–Pearson; likelihood-ratio test.)
[13] Cover, T. M., & Thomas, J. A. (2006). Elements of Information Theory (2nd ed.). Wiley. (Dataprocessing
inequality; Stein’s lemma; Kullback–Leibler divergence.)
[14] Coalition for Content Provenance and Authenticity (C2PA). (2026). C2PA Technical
Specification, v2.3. (Content Credentials; cryptographically signed provenance manifests.)
[15] Daubert v. Merrell Dow Pharmaceuticals, Inc., 509 U.S. 579 (1993). (Admissibility of expert
testimony; known or potential error rate.)
[16] Morrison, G. S. (2013). Tutorial on Logistic-Regression Calibration and Fusion: Converting a
Score to a Likelihood Ratio. Australian Journal of Forensic Sciences, 45(2), 173–197.
[17] European Network of Forensic Science Institutes (ENFSI). (2015). ENFSI Guideline for
Evaluative Reporting in Forensic Science (STEOFRAE), Approved Version 3.0.
[18] Goljan, M., Fridrich, J., & Chen, M. (2011). Defending Against Fingerprint-Copy Attack in
Sensor-Based Camera Identification. IEEE Transactions on Information Forensics and Security,
6(1), 227–236.
Forensic-Impact Articles
Mapping Threat Patterns Using Publicly Available Data
Guest Blogger: Ruqaya Osman Cybersecurity teams have long operated in two distinct lanes: those who investigate incidents after they occur, and those who gather intelligence to anticipate future threats. Digital Forensics and Incident Response (DFIR) practitioners...
Transition from Traditional Forensic Science to Digital Forensics: Challenges, Lessons, and Opportunities
Guest Blogger: Vaishnavi M.A. My journey into forensic science began with a desire to pursue something unique and intellectually stimulating. I have always been curious by nature and enjoyed solving problems, which led me to choose forensic science as my field of...
Investigative Perspective Claude Mythos AI
Guest Blogger: Sheetal Kumari At: MAD Forensics With the new age of AI, which has already arrived , showing continuous growth has introduced the world to a technology that is making advancements towards making human life easier yet also creating new dilemmas with...



