Article
What files can ChatGPT actually read?
Most tools accept almost anything. What changes between formats is how much of the file survives the trip — and whether the thing you get back was read or inferred.
Open a file PDF, Word, PowerPoint, PNG, JPG, WebP, Text · XLSX, CSV, Parquet, JSONPlain text and Markdown: nothing to go wrong
A .txt or .md file is already the thing a model reads. There is no conversion step, no layout to lose and no format to misjudge. If something can be saved as plain text, that is the version least likely to surprise you.
Word documents: usually clean, layout aside
A .docx carries its text, so extraction is reliable. What tends to go missing is the arrangement — footnotes, headers, tracked changes, comments and text sitting inside shapes or text boxes may or may not come through, and you will not be told which.
The older .doc format is a separate matter and support for it is patchy. Saving as .docx first costs nothing.
PDF: the format that splits in two
A PDF exported from an application contains letters and behaves like a Word file. A PDF produced by a scanner or a camera contains a picture of a page and contains no letters at all, which is why the same format can work perfectly one day and return nothing the next.
Nothing about the two looks different on screen. Press Ctrl+F and search for a word you can see: if the search fails, the file is a scan and needs OCR or a tool that reads images.
Spreadsheets: read fine, counted badly
An .xlsx opens and the rows come through. The catch is arithmetic. A language model that has been handed rows produces a total the way it produces a sentence — by prediction — and on a long sheet the answer drifts, quietly and plausibly.
Small tables are usually fine because the model can hold the whole thing at once. The files where a wrong total actually costs you something are exactly the ones too big for that.
CSV, TSV and JSON: shape matters more than size
These are plain text, so they read cleanly. What trips them up is structure: a CSV whose header is on row four, a semicolon-separated file from a European export, a JSON document nested six levels deep with the useful part on the fifth.
Saying what the file contains before asking about it saves more time than any format conversion.
Images: not a format question at all
PNG, JPG, WebP and AVIF all read, because the model is looking at the picture rather than parsing the file. The variable is not the extension but the photograph: focus, glare, angle and how much of the page is in frame.
HEIC, the format an iPhone saves by default, is the awkward one — many tools cannot decode it. Sharing the photo rather than sending the file usually hands you a JPG instead.
Where it stops
Audio and video are a different pipeline; a file that needs transcribing is not being read. Archives have to be unpacked first — a .zip is a box, not a document. Encrypted files, CAD drawings and proprietary formats from specialist software are generally out.
The honest summary: if the file is text, or a picture of text, it can be read. If it is neither, something has to convert it before anything can.
What CypherScan takes
PDF, Word and PowerPoint, images in PNG, JPG, WebP, AVIF and HEIC, plain text and Markdown, and the data formats: .xlsx, .xlsm, .xls, .csv, .tsv, .parquet, .json and .ndjson. A scanned PDF is not a special case here, because the pages are read as pictures of pages.
Data files are handled differently from documents, and deliberately: the rows never leave your machine. What is sent is a short description of the table, and the query that answers your question runs in your own browser against the full file — so the totals are calculated rather than predicted.
Can ChatGPT read JSON files?
Yes — JSON is plain text, so it goes in cleanly. Depth is the practical limit rather than size: heavily nested documents are easy to read and hard to reason about, so say which part matters.
Can ChatGPT read CSV files?
Yes, and CSV is often the most reliable way to hand over a table. Watch the separator — exports from European software frequently use semicolons — and check that the header really is the first row.
Can ChatGPT read DOCX files?
Yes. The text comes through reliably; footnotes, comments, tracked changes and anything inside a text box are the parts that may not, without any warning that they went missing.
Can it read a .parquet file?
Not generally, no — Parquet is a compressed columnar format rather than text, so most tools cannot open it at all. This site does, along with .xlsx and .ndjson.
Is there a file size limit?
Always, though it is rarely the number that bites. A long document is usually cut short rather than refused, which is worth knowing: an answer about the first forty pages of a two-hundred-page file looks exactly like an answer about the file.
What about .pages, .numbers and .key?
Apple's formats are not widely supported. Export to PDF, DOCX or CSV first — the export is lossless enough for anything you would want to ask about.