A MIME type, formally a media type, is a label such as image/png or application/pdf. Browsers use it to decide how to display content, and servers use it to decide what to accept. The label is only useful if it is true, and testing type validation is mostly a matter of finding out whose word the server takes.
Three sources, three levels of trust
When a file is uploaded, its type can be read from three places.
- The file name. The extension is whatever the user typed. It costs nothing to change.
- The Content-Type of the upload part. The browser fills this in, usually from the extension. A script can send anything.
- The content. Most binary formats begin with fixed signature bytes, often called magic bytes. Faking these takes real effort, and a file that passes a full parse is what it claims to be.
Sound validation uses the name and the header as hints for a helpful error message and makes the decision on content.
Signature bytes of common formats
| Format | MIME type | Leading bytes (hex) | As text |
|---|---|---|---|
application/pdf | 25 50 44 46 2D | %PDF- | |
| PNG | image/png | 89 50 4E 47 0D 0A 1A 0A | .PNG.... |
| JPEG | image/jpeg | FF D8 FF | |
| GIF | image/gif | 47 49 46 38 | GIF8 |
| WebP | image/webp | 52 49 46 46, then 57 45 42 50 at byte 8 | RIFF....WEBP |
| ZIP, DOCX, XLSX, PPTX, EPUB | application/zip and others | 50 4B 03 04 | PK.. |
| GZIP | application/gzip | 1F 8B | |
| MP4 | video/mp4 | 66 74 79 70 at byte 4 | ftyp |
| MP3 with ID3 tag | audio/mpeg | 49 44 33 | ID3 |
| WAV | audio/wav | 52 49 46 46, then 57 41 56 45 at byte 8 | RIFF....WAVE |
Every format page on this site lists the signature for that format, and you can look at the first bytes of any file yourself:
xxd -l 16 sample-png-10kb.png
file --mime-type -b sample-png-10kb.pngThe first command prints the leading 16 bytes in hex. The second asks the file utility to identify the type from content, and prints image/png.
Formats that cannot be recognised by their first bytes
Signature checks have limits, and good tests aim at them.
- ZIP-based formats. DOCX, XLSX, PPTX and EPUB all begin with the ZIP signature. Telling them apart means opening the archive and looking for
[Content_Types].xmlor themimetypeentry. - RIFF-based formats. WebP and WAV share their first four bytes. The distinguishing tag is at byte 8.
- Text formats. CSV, JSON, plain text and most configuration files have no signature. The only real check is to parse them.
- SVG and HTML. These are text, yet browsers treat them as active documents. Accepting them as "images" or "text" without sanitising is a common route to cross-site scripting.
Test cases
Use genuine sample files so that you know the correct answer in advance.
- Baseline. Upload one valid file of each accepted type. Confirm that the type your application stores matches the MIME type on the format page.
- Renamed file. Copy a PNG and give it a
.pdfextension. Upload it to a PDF-only field. Expected: refused. - Wrong header. Send a real PDF but declare it as
image/png, using the command below. Expected: the server decides from content, so it either accepts the file as a PDF or refuses it. It must never store a PDF labelled as a PNG. - No extension. Upload the file without an extension. Expected: handled by content, or refused with a clear message.
- Truncated file. Keep only the first 100 bytes of a PNG with
head -c 100. The signature is intact but the image is not. A signature-only check accepts it; a decoder does not. Decide which behaviour you need. - Office documents. Upload DOCX, XLSX and PPTX samples to a field that accepts them. Detection libraries frequently report
application/zip, and an allow-list that omits it rejects valid documents.
The wrong-header case can be sent with curl, which lets you set the part's type explicitly:
curl -sS -o /dev/null -w '%{http_code}\n' \
-F 'file=@sample-pdf-10kb.pdf;type=image/png' \
https://your-app.example/uploadA minimal content check
This Node.js function reads the first bytes of a file and matches them against a few signatures. Real projects should use a maintained detection library, but the logic is the same.
import { open } from 'node:fs/promises';
const SIGNATURES = [
{ mime: 'application/pdf', bytes: [0x25, 0x50, 0x44, 0x46, 0x2d] },
{ mime: 'image/png', bytes: [0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a] },
{ mime: 'image/jpeg', bytes: [0xff, 0xd8, 0xff] },
{ mime: 'image/gif', bytes: [0x47, 0x49, 0x46, 0x38] },
{ mime: 'application/zip', bytes: [0x50, 0x4b, 0x03, 0x04] },
];
export async function sniff(path) {
const handle = await open(path);
const { buffer } = await handle.read(Buffer.alloc(16), 0, 16, 0);
await handle.close();
const hit = SIGNATURES.find((s) => s.bytes.every((b, i) => buffer[i] === b));
return hit ? hit.mime : 'application/octet-stream';
}Serving files back
Validation on the way in is half of the job. When a stored file is downloaded or displayed:
- Send the
Content-Typeyou determined at upload, never one supplied by the uploader. - Send
X-Content-Type-Options: nosniff, which tells browsers not to second-guess the declared type. - For anything that is not meant to be displayed inline, send
Content-Disposition: attachment. - Serve user uploads from a separate domain where possible, so that a file that does execute has no access to your site's cookies.
To test this, upload an HTML sample, request its stored URL in a browser and confirm that it downloads instead of rendering.
Further reading
The rules browsers follow are defined in the WHATWG MIME Sniffing Standard, and registered types are listed in the IANA media types registry.