Sample files for MIME type and content validation testing

Files with correct, documented signatures for testing type detection and upload allow-lists.

There are three sources of truth about a file's type, and they can disagree: the extension in its name, the Content-Type the client declares, and the bytes of the file itself. Only the last is hard to fake.

Every format page here documents the MIME type and the leading signature bytes of that format, so you can check what your detection library reports against a file whose type is certain.

Recommended files

FileSizeWhy this oneDownload
10 KB PNG samplesample-png-10kb.png10 KB10,240 bytesUnambiguous eight-byte signatureDownload PNG
50 KB DOCX samplesample-docx-50kb.docx50 KB51,200 bytesZIP-based format that is easily misdetectedDownload DOCX
10 KB WebP samplesample-webp-10kb.webp10 KB10,240 bytesRIFF container shared with WAVDownload WebP
File without an extensionsample-file-without-extension61 B61 bytesNo extension to rely onDownload TXT
1 KB SVG samplesample-svg-1kb.svg1 KB1,024 bytesText file that browsers treat as an active documentDownload SVG

Checklist

  1. Verify detection on genuine files

    Run your detector over one sample of each accepted format and compare with the MIME type on the format page.

  2. Rename a file

    Change the extension of a PNG to .pdf and upload it. Validation based on content should refuse it.

  3. Lie in the header

    Send a real PDF with Content-Type: image/png. The server should trust neither the header nor the name.

  4. Test the ambiguous cases

    DOCX, XLSX, PPTX and EPUB all begin with the ZIP signature. Plain text, CSV and INI have no signature at all.

  5. Check what you serve back

    Stored files should be returned with the detected type and X-Content-Type-Options: nosniff.

Guides

Related use cases

Frequently asked questions

Is checking the file extension enough?
No. Anyone can rename a file. Use the extension as a first filter, then confirm the content with a signature check or a real parser.
Why does my library report application/zip for a Word document?
A DOCX file is a ZIP archive. Libraries that only read the leading bytes cannot tell the difference; better ones inspect the entries inside.