The file you recovered might be a ghost: checking integrity from the bytes alone
You run a recovery tool over a drive. It finishes with a number that feels like good news — 3,148 files found. Then you start opening them. Maybe half don't work. Some open and show a grey rectangle. Some are accepted by the viewer and display the wrong picture. Some are 4 MB on disk and contain nothing at all.
Recovery software isn't lying to you. It found the directory entry: the name, the size, the location the file used to occupy. What it can't know is whether the clusters behind that entry are still the clusters that belong to it.
Deleting a file doesn't erase it. On NTFS the entry is flagged and the space is marked free, and the bytes sit there until something else is written over them. Which means there are three independent things that can go wrong, and recovery tools report all three as "found".
The data was overwritten. The name came back; the contents are now somebody else's.
The data was partly overwritten. The first N clusters came back, the rest is gone. This is the nasty one — the file has a plausible size, a correct header, and opens in some viewers while rendering garbage.
The data never existed in the file you got. The tool knew the file was supposed to be 4 MB, so it wrote you 4 MB. The first two bytes and the last two bytes come from the format; everything between them is filler.
You can't ask the filesystem which case you're in. But you can ask the bytes.
Step 1: head and tail
Most formats put a fixed signature at the start and a terminator at the end. If either is missing, the file is cut.
| Format | Must start with | Must end with |
|---|---|---|
| JPEG | FF D8 FF | FF D9 (EOI) |
| PNG | 89 50 4E 47 0D 0A 1A 0A | IEND + CRC (49 45 4E 44 AE 42 60 82) |
%PDF- | %%EOF in the trailing couple of KB | |
| DOCX / XLSX / ZIP | PK | end-of-central-directory 50 4B 05 06 |
| MP4 / MOV | an ftyp box | a moov index, at either end |
Note the asymmetry, because getting this wrong marks valid files broken. A JPEG's EOI is the last two bytes, so you look at a 2-byte window. A ZIP's EOCD is somewhere in the last 64 KB + 22 bytes, because an archive comment can follow it. A PDF's %%EOF can be anywhere in the trailing 2 KB — and there can be more than one if the file was incrementally updated. Hardcode "check the last N bytes" across all formats and you'll produce false alarms all afternoon.
And then there's MP4, where the terminator isn't at the end at all. moov is a top-level box that can sit before or after mdat, so a tail window won't find it. You walk the box tree instead — each box is a 4-byte size followed by a 4-byte type, so you can hop from one to the next:
async function hasMoov(blob) {
let off = 0;
for (let i = 0; i < 64; i++) {
const h = await range(blob, off, 16);
if (h.length < 8) return false;
if (String.fromCharCode(h[4], h[5], h[6], h[7]) === 'moov') return true;
let sz = h[0] * 2 ** 24 + h[1] * 2 ** 16 + h[2] * 256 + h[3];
if (sz === 1) { sz = 0; for (let k = 8; k < 16; k++) sz = sz * 256 + h[k]; }
if (sz === 0 || sz < 8) return false; // 0 means "extends to end of file"
off += sz;
if (off >= blob.size) return false;
}
return false;
}
The cap on that loop is not decoration. A truncated file hands you a box size that points past the end of the data, and an unbounded walk will happily chase it forever.
Step 2: the hole in the middle
Here's the case that defeats every "is it valid" check: a file with a correct header, a correct terminator, and nothing in between.
Recovery tools routinely reconstruct a file to its declared size. If the clusters that held the middle are gone, the tool pads. You get a 4 MB JPEG whose first two bytes and last two bytes are perfect and whose contents are filler. A check that only looks at the head and the tail reports it as fine.
So you sample the middle. And the obvious way to do that — count zero bytes — has a hole in it big enough to drive a recovery job through.
Not every tool pads with zero. Some pad with 0xFF. Some repeat a short pattern. If your detector is if (zeroRatio > 0.98), a file padded with 0xFF sails straight through it, and those files are not rare.
The fix is to stop looking for a specific value and look for the absence of variety instead. A 8 KB chunk of real compressed data uses essentially all 256 byte values. A chunk of padding uses one, maybe four. That distinction catches every padding style at once:
function analyze(u8) {
const seen = new Uint8Array(256);
let zero = 0;
for (let i = 0; i < u8.length; i++) { const b = u8[i]; if (b === 0) zero++; seen[b] = 1; }
let distinct = 0;
for (let i = 0; i < 256; i++) if (seen[i]) distinct++;
return { zeroRatio: +(zero / u8.length).toFixed(3), distinct };
}
const isHole = (a) => a.zeroRatio > 0.98 || (a.distinct > 0 && a.distinct <= 4);
Keep the zero test as well — it earns its place, because a chunk that is 99% zeros with a few surviving bytes is still a hole, and 99% zeros will show up as, say, 30 distinct values.
Step 3: sample proportionally
Three evenly spaced points catch a contiguous hole, which is what overwriting usually produces. But three points on a 2 GB video will miss a 6 MB gap, and 6 MB is a lot of missing video.
So scale the count with the size and cap it, because the whole point of sampling is not reading the file:
function sampleOffsets(size) {
if (size < 128 * 1024) return [];
const n = Math.min(32, Math.max(3, Math.ceil(size / (4 * 1024 * 1024))));
const step = size / (n + 1);
const pts = [];
for (let i = 1; i <= n; i++) pts.push(Math.floor(i * step));
return pts;
}
Small files return nothing on purpose. Below about 128 KB the file is all edge and no middle — sampling it produces noise, not signal, and a thumbnails-sized JPEG would get flagged for having a boring interior.
Where this deliberately stops
Uncompressed formats are exempt. A real black BMP is genuinely, correctly almost all zeros. So is a minute of digital silence in a WAV. If you run a hole detector over those you will condemn perfectly good files, so check the format first and skip the sampling when it isn't compressed. This is the one place where knowing the type before judging the contents matters.
Structure is not content. A JPEG with an SOI, an EOI and no holes can still contain the wrong photograph, because the clusters it was rebuilt from belonged to a different file that happened to sit nearby. No byte-level check can detect that. Only a human opening it can.
Damaged doesn't mean worthless. A file we call partial still carries whatever fraction did survive. Photo recovery in particular: a JPEG missing its tail often still renders most of the image, and if it's the only copy of a photograph, "most of the image" is a very good outcome. Don't let a checker talk you into deleting anything.
And one filesystem note, because it changes your odds before you start: ext4 clears block pointers on delete, where NTFS mostly leaves them. Filenames on ext4 are usually lost outright, which is why Linux recovery is harder than Windows recovery even when the data itself survived. If you're working an ext4 volume, do it from a live USB — never from the installed system, which is writing to the same disk you're trying to rescue.
The whole thing, ready to paste
Assembled. It takes a File (or a Blob), reads only the windows it needs via Blob.slice() — which is lazy, so nothing is read until you ask for the bytes — and returns a verdict with the evidence behind it.
Paste it into a DevTools console, then either call it on a file handle you already have, or set up a drop target and drag a folder in:
document.addEventListener('dragover', (e) => e.preventDefault());
document.addEventListener('drop', async (e) => {
e.preventDefault();
console.table(await ghostCheck.bulk(e.dataTransfer.files));
});
const ghostCheck = (() => {
const KB = 1024, MB = 1024 * 1024;
const CHUNK = 8 * KB;
const MIN_SAMPLABLE = 128 * KB;
// Blob.slice() is lazy: nothing is read until you ask for the bytes.
async function range(blob, start, len) {
const s = Math.max(0, Math.round(start));
const e = Math.min(blob.size, s + len);
if (e <= s) return new Uint8Array(0);
return new Uint8Array(await blob.slice(s, e).arrayBuffer());
}
function startsWith(u8, sig) {
if (u8.length < sig.length) return false;
for (let i = 0; i < sig.length; i++) if (u8[i] !== sig[i]) return false;
return true;
}
function indexOf(u8, sig) {
outer: for (let i = 0; i + sig.length <= u8.length; i++) {
for (let j = 0; j < sig.length; j++) if (u8[i + j] !== sig[j]) continue outer;
return i;
}
return -1;
}
const ascii = (s) => [...s].map((c) => c.charCodeAt(0));
function analyze(u8) {
const seen = new Uint8Array(256);
let zero = 0;
for (let i = 0; i < u8.length; i++) { const b = u8[i]; if (b === 0) zero++; seen[b] = 1; }
let distinct = 0;
for (let i = 0; i < 256; i++) if (seen[i]) distinct++;
return { zeroRatio: +(zero / u8.length).toFixed(3), distinct };
}
// A hole is not always zeros. Recovery tools pad with 0x00, sometimes 0xFF,
// sometimes a short repeating pattern. Both look the same from here:
// a chunk of real compressed data uses ~256 distinct byte values; padding uses ~1.
const isHole = (a) => a.zeroRatio > 0.98 || (a.distinct > 0 && a.distinct <= 4);
function sampleOffsets(size) {
if (size < MIN_SAMPLABLE) return [];
const n = Math.min(32, Math.max(3, Math.ceil(size / (4 * MB))));
const step = size / (n + 1);
const pts = [];
for (let i = 1; i <= n; i++) pts.push(Math.floor(i * step));
return pts;
}
const FORMATS = [
{ id: 'jpeg', label: 'JPEG', head: [0xff, 0xd8, 0xff], tail: [0xff, 0xd9], tailWindow: 2, compressed: true },
{ id: 'png', label: 'PNG', head: [0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a], tail: [0x49, 0x45, 0x4e, 0x44, 0xae, 0x42, 0x60, 0x82], tailWindow: 16, compressed: true },
{ id: 'zip', label: 'ZIP / Office', head: [0x50, 0x4b], tail: [0x50, 0x4b, 0x05, 0x06], tailWindow: 65557, compressed: true },
{ id: 'pdf', label: 'PDF', head: ascii('%PDF-'), tailText: '%%EOF', tailWindow: 2048, compressed: true },
{ id: 'bmp', label: 'BMP', head: ascii('BM'), compressed: false },
{ id: 'wav', label: 'WAV', head: ascii('RIFF'), compressed: false },
];
const FTYP = ascii('ftyp');
// Walk top-level ISO-BMFF boxes looking for `moov`. Capped: a truncated file
// can point the offset past the end, and stopping beats looping.
async function hasMoov(blob) {
let off = 0;
for (let i = 0; i < 64; i++) {
const h = await range(blob, off, 16);
if (h.length < 8) return false;
if (String.fromCharCode(h[4], h[5], h[6], h[7]) === 'moov') return true;
let sz = h[0] * 2 ** 24 + h[1] * 2 ** 16 + h[2] * 256 + h[3];
if (sz === 1) { // 64-bit extended size
sz = 0;
for (let k = 8; k < 16; k++) sz = sz * 256 + h[k];
}
if (sz === 0 || sz < 8) return false; // 0 means "to end of file"
off += sz;
if (off >= blob.size) return false;
}
return false;
}
async function check(blob) {
const size = blob.size;
let bytesRead = 0;
const headBytes = await range(blob, 0, 16); bytesRead += headBytes.length;
const fmt = FORMATS.find((f) => startsWith(headBytes, f.head)) || null;
let type = fmt ? fmt.label : 'unknown';
const compressed = fmt ? fmt.compressed : true;
if (!fmt && startsWith(headBytes.subarray(4, 8), FTYP)) type = 'MP4 / MOV';
let tail = null; // null = this format has no terminator worth checking
if (type === 'MP4 / MOV') {
tail = await hasMoov(blob);
} else if (fmt && (fmt.tail || fmt.tailText)) {
const sig = fmt.tail || ascii(fmt.tailText);
const win = await range(blob, size - fmt.tailWindow, fmt.tailWindow);
bytesRead += win.length;
tail = indexOf(win, sig) >= 0;
}
const holes = [];
let minDistinct = null;
if (compressed) {
for (const at of sampleOffsets(size)) {
const c = await range(blob, at, CHUNK); bytesRead += c.length;
if (c.length < 512) continue;
const a = analyze(c);
if (minDistinct === null || a.distinct < minDistinct) minDistinct = a.distinct;
if (isHole(a)) holes.push({ at, ...a });
}
}
const cut = tail === false;
const holey = holes.length > 0;
let verdict;
if (type === 'unknown') verdict = 'UNRECOGNISED';
else if (!cut && !holey) verdict = tail === null ? 'PLAUSIBLE (no terminator)' : 'INTACT';
else if (!cut && holey) verdict = 'GHOST';
else if (cut && holey) verdict = 'GHOST + PARTIAL';
else verdict = 'PARTIAL';
return {
type, size, tail: tail === null ? '—' : tail, holes: holes.length,
minDistinct: minDistinct === null ? '—' : minDistinct,
kbRead: +(bytesRead / 1024).toFixed(2), verdict,
};
}
check.bulk = async (files) => Promise.all(
[...files].map(async (f) => ({ file: f.name, ...(await check(f)) }))
);
return check;
})();
Against synthetic files — a complete JPEG, one cut before its EOI, a complete PNG, a DOCX with no EOCD, a JPEG with a single hole, a padded JPEG, the same padded with 0xFF instead of zero, a legitimate black BMP, a complete and a truncated MP4, and a 40 KB thumbnail:
┌─────────┬──────────────────────┬────────────────┬───────┬───────┬─────────────┬────────┬─────────────────────────────┐
│ (index) │ file │ type │ tail │ holes │ minDistinct │ kbRead │ verdict │
├─────────┼──────────────────────┼────────────────┼───────┼───────┼─────────────┼────────┼─────────────────────────────┤
│ 0 │ 'photo-complete.jpg' │ 'JPEG' │ true │ 0 │ 256 │ 24.02 │ 'INTACT' │
│ 1 │ 'photo-cut.jpg' │ 'JPEG' │ false │ 0 │ 256 │ 24.02 │ 'PARTIAL' │
│ 2 │ 'scan.png' │ 'PNG' │ true │ 0 │ 256 │ 24.03 │ 'INTACT' │
│ 3 │ 'report.docx' │ 'ZIP / Office' │ false │ 0 │ 256 │ 88.04 │ 'PARTIAL' │
│ 4 │ 'photo-hole.jpg' │ 'JPEG' │ true │ 2 │ 1 │ 24.02 │ 'GHOST' │
│ 5 │ 'photo-ghost.jpg' │ 'JPEG' │ true │ 3 │ 1 │ 24.02 │ 'GHOST' │
│ 6 │ 'photo-ghost-ff.jpg' │ 'JPEG' │ true │ 3 │ 1 │ 24.02 │ 'GHOST' │
│ 7 │ 'black.bmp' │ 'BMP' │ '—' │ 0 │ '—' │ 0.02 │ 'PLAUSIBLE (no terminator)' │
│ 8 │ 'clip.mp4' │ 'MP4 / MOV' │ true │ 0 │ 256 │ 24.02 │ 'INTACT' │
│ 9 │ 'clip-truncated.mp4' │ 'MP4 / MOV' │ false │ 0 │ 256 │ 24.02 │ 'PARTIAL' │
│ 10 │ 'thumb-small.jpg' │ 'JPEG' │ true │ 0 │ '—' │ 0.02 │ 'INTACT' │
└─────────┴──────────────────────┴────────────────┴───────┴───────┴─────────────┴────────┴─────────────────────────────┘
Three rows in there are the ones worth the effort.
Row 6 is the 0xFF-padded file. Head intact, tail intact, not a single zero byte in it. A zero-counting detector scores it perfectly clean.
Row 7 is a real black BMP, exempt from hole detection because it's uncompressed. It's almost entirely zeros and it is a perfectly good file.
Row 3 cost 88 KB to judge a 500 KB file, because the ZIP end-of-central-directory can hide anywhere in the last 64 KB. That's the price of per-format windows — and it's still a fifth of the file, rather than the whole thing.
Practical notes
Run this before you open recovered files, not after. Opening a damaged file is usually harmless, but opening it with an application that writes — a photo editor generating thumbnails, a database creating a lock file — writes to the same drive you're recovering from.
Compare the size the recovery tool reported against the size you actually got. If the tool said 4 MB and the file on disk is 900 KB, you don't need any of the above.
Sort by verdict before you sort by anything else, and copy the intact ones off first. On a drive that's actively failing, the order you work in is worth more than the tool you picked.
How this was written
Disclosure first: the draft of this post was written by an AI agent from a problem I chose, with one standing instruction — no claim gets written down until it has been executed. The eleven test cases in the table above were built as byte arrays, fed through the function, and their output pasted here. None of it is illustrative.
I'll say what that bought, because it's the only part of this process worth reading about.
The original plan was to detect a hollow file the obvious way: count zero bytes in the middle and call it empty. Then I built a fake JPEG padded with 0xFF instead of 0x00 — head intact, tail intact, two megabytes of nothing, and not a single zero byte anywhere. The naive check reported it as a healthy file. So did every "verify your recovered files" snippet I could find, because they all ask one question and the question is wrong.
That produced the actual rule in this post: ask whether a chunk of the file has variety, not whether it has zeros. Real compressed data will touch all 256 byte values in an 8 KB sample. Padding touches one. The reframed question catches zero-fill, 0xFF-fill and short repeating patterns with the same three lines of code.
Which is a specific, small version of a general thing: the tests were worth more than the writing. A model writes a confident paragraph about either approach in the same number of seconds. Only the run tells you which one is wrong.