New Document Converter Detects True Format Beyond Misleading File Extensions and Headers

October 3, 2026
New Document Converter Detects True Format Beyond Misleading File Extensions and Headers
  • The author implemented automatic detection for a universal document converter and found that file extensions and Content-Type headers can be misleading when determining a document’s true format.

  • DOCX, PPTX, XLSX, and EPUB are all ZIP-based archives that share the PK signature, so correct format must be determined from within the archive, not just by filename or header.

  • Originating from a Moltbook AI agent post, the concept is operationalized via an endpoint hosted at the provided URL for converting documents to Markdown.

  • Plain HTML and plain text lack an archive structure, so detection relies on conservative heuristics that avoid guessing and tend to fail loudly when uncertain.

  • Two main lessons emerge: never trust the filename alone and only partially trust the header, since proxies can alter one piece of evidence; for shared signatures, opening the archive and reading its manifest yields the correct identification.

  • A file named contract.pdf was actually a DOCX because a proxy stamped the header based on the filename, while the container contents differed from the extension.

  • Office formats place a [Content_Types].xml at the archive root and use paths like word/, ppt/, xl/ to indicate the precise format inside the ZIP container.

  • EPUB distinguishes itself by including a mimetype file as the first ZIP entry containing application/epub+zip, stored uncompressed for immediate recognition.

  • The ideas have been consolidated into an auto-detecting endpoint that accepts bytes or a URL and outputs Markdown, combining multiple converters into a single tool.

Summary based on 1 source


Get a daily email with more Tech stories

More Stories