Your file names aren't corrupt. They're just being read in the wrong alphabet.
You open a ZIP that came from a colleague and this is what's inside:
الصورة.jpg
تقرير نهائي.pdf
Işık.txt
The files open fine. The photo is intact, the PDF renders. Only the names are unreadable, and when there are four hundred of them, four hundred unreadable names is a genuine problem — you can't sort them, search them, or hand them back to the person who asked.
If you've mostly worked with English filenames, this looks like corruption. It isn't. Nothing happened to the bytes. Every one of those names was read with the wrong character encoding, which is a much better situation to be in, because the fix is arithmetic rather than guesswork.
What actually happened
الصورة is six Arabic letters. In UTF-8 each one takes two bytes, so the name is twelve bytes:
D8 A7 D9 84 D8 B5 D9 88 D8 B1 D8 A9
ا ل ص و ر ة
Now read those exact twelve bytes one at a time, treating each byte as a whole character in a single-byte code page — Windows-1252, say:
D8 A7 D9 84 D8 B5 D9 88 D8 B1 D8 A9
Ø § Ù „ Ø µ Ù ˆ Ø ± Ø ©
Put them together and you get الصورة. The bytes crossed the wire perfectly. They were just decoded against the wrong table.
That's the entire accident. It's called mojibake, and it is by far the most common reason a "damaged" filename is actually fine.
The shape of the mess tells you which table was used
| Read as | What it looks like | Where it comes from |
|---|---|---|
| Windows-1256 (Arabic) | ط§ظ„طµظˆط±ط© — still Arabic letters, but nonsense | An Arabic Windows box reading UTF-8 bytes |
| Windows-1252 / Latin-1 | الصورة — accented Latin, stray symbols | ZIP archives, FTP, cross-system copies |
| Windows-1254 (Turkish) | Işık for Işık | Turkish Windows, older archives |
| Mac Arabic | ظ\'عÑ — Arabic mixed with ASCII punctuation | Archives created on old Macs |
| DOS Arabic (cp864) | ﻅ§ﻋ▒ — box drawing and filler glyphs | Very old DOS-era archives |
| Applied twice | ال — a mess with a mess inside it | A name that was already wrong, converted again |
That last row matters more than it looks. Double-encoded names are common: somebody sees the mess, runs it through a converter to "fix" it, and produces a deeper mess. The operation is reversible in reverse, so the fix is just to apply it twice.
Doing it in the browser
Decoding is free. TextDecoder implements the WHATWG encoding standard:
const bytes = new Uint8Array([0xd8, 0xa7, 0xd9, 0x84]);
new TextDecoder('windows-1252').decode(bytes); // 'ال'
new TextDecoder('windows-1256').decode(bytes); // 'ط§ظ„'
But there's a limit worth knowing before you build anything on this. I tried every Arabic-relevant label in Node 22:
windows-1256 OK
windows-1252 OK
windows-1254 OK
x-mac-arabic The "x-mac-arabic" encoding is not supported
ibm864 The "ibm864" encoding is not supported
Three out of five. WHATWG's list includes x-mac-cyrillic but no x-mac-arabic, and no DOS code pages at all. So the two encodings behind the oldest archives — the ones most likely to be mangled in the first place — are exactly the ones TextDecoder won't touch. For those you need your own 256-entry table, dumped from Python (bytes.decode('mac_arabic')) or from the Unicode Consortium's mapping files, shipped as JSON.
Then comes the part that trips everyone: going backwards means char → byte, and the browser won't do that for you.
new TextEncoder().encode('Ø'); // [0xc3, 0x98] ← UTF-8, always
TextEncoder is hardcoded to UTF-8. There is no new TextEncoder('windows-1252') and there won't be one. So you build the reverse table yourself, which takes ten lines:
function buildEncoder(label) {
const dec = new TextDecoder(label);
const table = new Map();
for (let b = 0; b < 256; b++) {
const ch = dec.decode(new Uint8Array([b]));
if (ch.length !== 1) continue;
if (ch === '\uFFFD') continue; // undefined slot — several bytes land here
if (!table.has(ch)) table.set(ch, b);
}
return table;
}
The '\uFFFD' guard is not cosmetic. Legacy code pages have holes — Windows-1252 leaves 0x81, 0x8D, 0x8F, 0x90 and 0x9D undefined — and the decoder maps all of them to U+FFFD. Skip that check and five different bytes collapse onto one character in your reverse table, and you corrupt data on the way back without any error.
Now the repair itself. One pass, with a detail most implementations miss:
function pass(s, label) {
const t = buildEncoder(label);
const bytes = [];
for (const ch of s) {
const cp = ch.codePointAt(0);
if (cp < 0x80) { bytes.push(cp); continue; } // ASCII survived the trip
const b = t.get(ch);
if (b === undefined) return null; // not reversible in this page
bytes.push(b);
}
try {
// fatal:true is the filter most implementations forget. A wrong code page
// usually yields bytes that aren't valid UTF-8 — this rejects them for free.
return new TextDecoder('utf-8', { fatal: true }).decode(new Uint8Array(bytes));
} catch { return null; }
}
That fatal: true is doing real work. Without it, TextDecoder inserts U+FFFD for invalid sequences and hands you back a string that looks like a result. With it, wrong candidates throw and get dropped — one less thing your ranking has to sort out later.
When it genuinely cannot be fixed
Two cases are not recoverable, and a tool that pretends otherwise is worse than useless.
The name contains replacement characters. If you see الص�رة — the diamond, or a plain ? where a letter should be — then at some point a decoder met bytes it couldn't represent and substituted. The original byte values are gone. There is no table that brings them back.
The name was truncated. Mojibake preserves length. If an 8.3 conversion or a filesystem limit chopped the name, there's nothing to decode.
Saying "this one is gone" beats returning a plausible-looking guess, because a wrong name gets pasted onto a file and from then on nobody can tell it was ever wrong.
Picking the right candidate
The awkward part: several code pages produce the same mojibake. Arabic bytes read through Windows-1252 and Windows-1254 come out identical, because those two pages share the range where Arabic UTF-8 bytes land. You will regularly get multiple structurally-valid decodings and have to choose between them.
Score them by where the characters sit in Unicode. No dictionary required:
const BLOCKS = {
arabic: [[0x0600, 0x06ff], [0x0750, 0x077f], [0xfb50, 0xfdff], [0xfe70, 0xfeff]],
latin: [[0x0041, 0x007a], [0x00c0, 0x017f], [0x0100, 0x017f]],
cyrillic: [[0x0400, 0x04ff]],
};
function score(s) {
let counted = 0;
const hits = {};
for (const ch of s) {
const c = ch.codePointAt(0);
if (c < 0x80) continue; // ASCII is neutral: extensions, digits, hyphens
counted++;
for (const [name, ranges] of Object.entries(BLOCKS)) {
if (ranges.some(([lo, hi]) => c >= lo && c <= hi)) hits[name] = (hits[name] || 0) + 1;
}
}
if (!counted) return 0;
return Math.max(0, ...Object.values(hits)) / counted;
}
Keep ASCII out of the denominator. I got this wrong on the first pass: scoring .jpg and -01 as characters makes every تقرير نهائي.pdf top out around 0.77, so you can never use a clean threshold. Extensions and digits are language-neutral — exclude them and real results land on 1.0.
This is also why "just try the most common encoding" doesn't work. Which page is correct depends on the language inside the name, which is precisely what you couldn't read a moment ago.
The whole thing, ready to paste
Assembled, with caching and up to three passes for double- and triple-encoded names. Paste it into a DevTools console and call fix('الصورة'):
const fix = (() => {
const LABELS = ['windows-1256', 'windows-1252', 'windows-1254'];
const cache = new Map();
function encoder(label) {
if (!cache.has(label)) {
const dec = new TextDecoder(label);
const t = new Map();
for (let b = 0; b < 256; b++) {
const ch = dec.decode(new Uint8Array([b]));
if (ch.length === 1 && ch !== '\uFFFD' && !t.has(ch)) t.set(ch, b);
}
cache.set(label, t);
}
return cache.get(label);
}
function pass(s, label) {
const t = encoder(label);
const bytes = [];
for (const ch of s) {
const cp = ch.codePointAt(0);
if (cp < 0x80) { bytes.push(cp); continue; }
const b = t.get(ch);
if (b === undefined) return null;
bytes.push(b);
}
try {
return new TextDecoder('utf-8', { fatal: true }).decode(new Uint8Array(bytes));
} catch { return null; }
}
const BLOCKS = {
arabic: [[0x0600, 0x06ff], [0x0750, 0x077f], [0xfb50, 0xfdff], [0xfe70, 0xfeff]],
latin: [[0x0041, 0x007a], [0x00c0, 0x017f], [0x0100, 0x017f]],
cyrillic: [[0x0400, 0x04ff]],
};
function score(s) {
let counted = 0;
const hits = {};
for (const ch of s) {
const c = ch.codePointAt(0);
if (c < 0x80) continue;
counted++;
for (const [name, ranges] of Object.entries(BLOCKS)) {
if (ranges.some(([lo, hi]) => c >= lo && c <= hi)) hits[name] = (hits[name] || 0) + 1;
}
}
if (!counted) return 0;
return +(Math.max(0, ...Object.values(hits)) / counted).toFixed(3);
}
return function fix(mangled, maxPasses = 3) {
const seen = new Map();
seen.set(mangled, { via: 'unchanged', score: score(mangled) });
for (const label of LABELS) {
let cur = mangled;
for (let i = 1; i <= maxPasses; i++) {
const next = pass(cur, label);
if (next === null || next === cur) break;
cur = next;
if (!seen.has(cur)) seen.set(cur, { via: `${label} x${i}`, score: score(cur) });
}
}
return [...seen.entries()]
.map(([text, m]) => ({ text, via: m.via, score: m.score }))
.sort((a, b) => b.score - a.score);
};
})();
fix('الصورة') — the double-encoded case — returns:
┌─────────┬─────────────────────────────┬───────────────────┬───────┐
│ (index) │ text │ via │ score │
├─────────┼─────────────────────────────┼───────────────────┼───────┤
│ 0 │ 'الصورة' │ 'windows-1252 x2' │ 1 │
│ 1 │ 'الصورة' │ 'unchanged' │ 0.52 │
│ 2 │ 'الصورة' │ 'windows-1252 x1' │ 0.5 │
└─────────┴─────────────────────────────┴───────────────────┴───────┘
And a name that already contains U+FFFD returns nothing usable — the honest outcome:
┌─────────┬───────────────┬─────────────┬───────┐
│ (index) │ text │ via │ score │
├─────────┼───────────────┼─────────────┼───────┤
│ 0 │ 'الص�رة' │ 'unchanged' │ 0.455 │
└─────────┴───────────────┴─────────────┴───────┘
Two production notes if you ship this. Cache the reverse tables — rebuilding three 256-entry maps per filename is wasteful when you have 400 of them. And always keep the unchanged row: a name that was already fine has nothing to reverse, and returning an empty result makes a working tool look broken.
Why bother
If your users are in the Gulf, Iran or Türkiye, this isn't an edge case. It happens every time an archive crosses a Windows box, every time something legacy touches a modern filename, every time a name survives a transfer but its encoding declaration doesn't. And the failure is quiet — nobody files a bug about a file whose name is merely ugly. They rename it, or they live with it.
The repair is deterministic, takes milliseconds, and runs entirely on the user's device. Reading a filename is not the same as reading the file, and most people would rather you didn't read the file.
How this was written
Here is the honest version: the draft was written by an AI agent, working from a problem I picked and under one rule I insisted on — nothing goes into the post until it has been run. Every snippet below was executed, and the tables are pasted from real output rather than from memory.
That rule did most of the work. Two claims I started with quietly failed on contact:
TextDecoderin Node 22 throws onx-mac-arabicandibm864. I assumed they'd work, because the WHATWG encoding list looks comprehensive — it isn't. It carriesx-mac-cyrillicand nothing else from the Mac Arabic family, and no DOS code pages at all.- The first version of the snippet returned an empty table for a filename that was never broken. A correct tool, reporting failure, for the most common input it would ever see.
Neither would have survived as a description. Both fell out of running the code. Which is the whole argument for running your examples — and for treating "the model says so" as a hypothesis, including when the model is the one writing your blog post.