About Integrity

A forensic tool built to surface the truth in every pixel.

Integrity is developed by the Computer Society of India, VIT Student Branch as an open-source project. It operates as a two-tier forensic pipeline, giving every uploaded image an instant verdict while reserving a heavier, more rigorous analysis for the cases that need it.

The quick scan tier is built for production speed: it strips out heavy dependencies like PyTorch and Torchvision, relying purely on NumPy and ONNX Runtime to fuse a spatial CNN, an FFT-based frequency analysis, and metadata scanning into one softmax prediction. When an image proves more difficult, heavily compressed, screenshotted, or generated by newer diffusion models such as Midjourney v6, the deep scan tier engages a heavier tri-stream pipeline built on PyTorch, OpenCV, and Pillow, adding an error-level-analysis branch that replaces the spatial branch entirely.

Computer mouse patent illustration

tier 1 — quick scan

Optimized to process incoming requests instantly and catch obvious AI-generated images by checking the metadata attached to them, such as C2PA cryptographic signatures, before spending heavier compute. Runs on a dual-stream + metadata architecture.

01

spatial branch — IntegritySpatialCNN

A custom-built CNN with squeeze-and-excitation attention blocks that focus on local generative anomalies. Image resizing and normalization are reimplemented in NumPy so the inference script can stay free of heavy frameworks. Outputs a 128-dimensional feature vector.

02

frequency branch — FFT center crop

Converts a 256×256 grayscale version of the image into the frequency domain with a Fast Fourier Transform, then isolates a 32×32 low/mid frequency center crop, a trap tuned to catch the repeating grid-like artifacts left behind by AI upscalers. Outputs a 64-dimensional vector.

03

metadata branch — hard gate

Performs a fast byte-level scan of the raw image for metadata attached to it, including standard EXIF data and C2PA cryptographic signatures, producing a simple 2-flag array (0.0 or 1.0).

04

fusion & prediction

The three streams are concatenated into a 194-dimensional vector (128 spatial + 64 frequency + 2 metadata). An ONNX Runtime model applies a softmax to produce the final real / fake probability, built for instant production response.

tier 2 — deep scan

A heavier, more rigorous pipeline that steps in when better spatial rendering or screenshot recompression blinds the quick scan's frequency trap. Runs on a tri-stream spatial + frequency + forensic ELA architecture.

01

spatial branch

The same deep CNN with squeeze-and-excitation blocks used in the quick scan, but condensed to a 64-dimensional vector to leave room for additional forensic signals.

02

frequency branch — high-frequency map

Generates a 1024-value high-frequency magnitude map via FFT, deliberately zeroing out the center 32×32 region to ignore low frequencies and isolate the microscopic, high-frequency noise typical of advanced diffusion models. Outputs a 64-dimensional vector.

03

forensic branch — error level analysis

Re-encodes the image at 90% JPEG quality entirely in memory, computes the pixel-level difference against the original, and amplifies it into an error-level-analysis map. AI-generated pixels compress differently from natural sensor noise, exposing artifacts hidden beneath the surface. This branch carries double voting weight, 128 dimensions, so it can override the spatial branch when needed.

04

fusion & prediction

Spatial (64), frequency (64), and forensic ELA (128) streams are fused into a 256-dimensional vector, passed through heavily regularized dropout layers to produce a high-confidence final verdict.

architecture

The quick scan tier runs entirely from quickscaninference.py, NumPy and ONNX Runtime only, with no PyTorch or Torchvision, so it can return a verdict the moment a request arrives. This keeps the runtime container small and the cold-start time low, which matters when the tool is fielding a steady stream of incoming uploads. The deep scan tier runs from deepscaninference.py, introducing PyTorch, OpenCV, and Pillow (ImageChops) for its error-level-analysis forensic step. Because it loads a full deep learning stack, it trades some of that speed for the ability to catch artifacts the quick scan's lighter pipeline is built to skip past.

Deep scan is the better choice for heavily compressed images, screenshots of AI-generated content rather than source files, and outputs from newer diffusion models, cases where stronger spatial rendering can blind the quick scan's frequency trap. The two tiers are designed to complement rather than replace each other: most uploads only need the quick scan's instant verdict, and the deep scan is reserved for the harder cases where that verdict comes back uncertain or where the stakes of getting it wrong are higher.

Both tiers return a confidence percentage and a categorical verdict of AI Generated, Likely Real, or Inconclusive, so the output stays consistent regardless of which pipeline produced it. That consistency is what lets the frontend treat the two tiers interchangeably, swapping in the heavier model only when the situation actually calls for it.