Your next AI task might need the words in a report, the figures in a budget table, or the instructions in a scanned document. Sending the original file can also ask the model to process the appearance of every page.
That distinction matters when you pay by the token.
SourceShelf converts PDFs, modern Office documents, and text in images into Markdown locally on your Mac. You can review the result, organize it into a focused Pack, and send the content your task needs to GPT-6 Astra.
We ran seven synthetic files through SourceShelf’s conversion code to explore the difference. Our local cost model estimated 84% fewer input tokens for a 40-page text PDF and 97% fewer for a sparse, 20-page scanned packet when comparing Markdown with a high-detail page-image scenario.
Those are illustrative estimates, not measured Astra invoices or a promise for every document. The biggest opportunity is removing page-image overhead from text-focused work. Markdown is not inherently smaller than plain text, and selecting an excerpt is a different kind of saving from converting a whole document.
Why the original format matters
OpenAI’s Responses API handles file types differently. PDFs supply both extracted text and page images to vision-capable models. Word and PowerPoint files supply extracted text, while spreadsheets follow a separate process that adds summaries and metadata. For Astra, PDF auto detail currently uses high. OpenAI file-input documentation
If your question concerns the wording of a policy, paying for page images may add little value. Converting that policy into reviewed Markdown gives you a way to send its text without those images.
For a chart, engineering drawing, or photograph, the visual information may be essential. Keep the relevant image in the request when the task depends on it.
GPT-6 Astra API pricing
As checked on September 20, 2026, GPT-6 Astra’s standard short-context API rates are:
| Token category | USD per million tokens |
|---|---|
| Input | $10.00 |
| Cached input | $1.00 |
| Cache writes | $12.50 |
| Output | $50.00 |
The calculations below use the $10 input rate to make the document representations comparable. They exclude output, tools, cache writes, and other request content. They are input-rate equivalents rather than predictions of a complete bill. OpenAI API pricing
At that rate, removing 100,000 input tokens is worth $1 per uncached read. Reusing compact source material across many independent tasks can make that difference meaningful.
Our local experiment
We created a 40-page report, a 20-page image-only packet, two text images, a Word report, a 30-slide PowerPoint deck, and a 500-row Excel workbook. All content was synthetic. The scanned packet was deliberately sparse; the report and second image contained much denser text.
We ran the files through the repository’s SourceShelf extractors, semantic Markdown renderer, and Markdown writer on a Mac. Counts include the generated source metadata and headings. The documents were converted, not summarized.
Text was counted with tiktoken 0.14.0 using o200k_base. That tokenizer does not explicitly map Astra, so these are local proxy counts, not an authoritative count of Astra’s text tokens. Image tokens were calculated using Astra’s documented patch rules. For PDFs, we assumed page images at 144 DPI and added locally extracted native text. OpenAI’s actual PDF rasterization and text extraction may differ. No files were sent to the API, and no API usage was billed during this experiment.
PDFs and text images: the largest opportunity
| Test document | Estimated original input tokens | Markdown proxy tokens | Estimated reduction | Estimated input cost (USD): original → Markdown |
|---|---|---|---|---|
| 40-page text PDF | 112,449 | 18,184 | 83.8% | $1.1245 → $0.1818 |
| 20-page scanned packet | 45,600 | 1,559 | 96.6% | $0.4560 → $0.0156 |
| Sparse text image, 1200 × 1600 | 2,280 | 162 | 92.9% | $0.0228 → $0.0016 |
| Dense text image, 1200 × 1600 | 2,280 | 545 | 76.1% | $0.0228 → $0.0055 |
These comparisons use high image detail. The chart normalizes each original to 100%; bar lengths compare representations within a document, not token totals across documents. For example, a 1200 × 1600 image occupies 38 × 50 patches of 32 pixels, or 1,900 patches. Astra’s 1.2 multiplier produces 2,280 image tokens. Turning the sparse image into 162 tokens of reviewed Markdown removes most of that input cost. OpenAI image token calculations
For the 40-page report, the modeled saving is about $0.94 per uncached read, or $94.27 across 100 independent reads at the same rate. That is an extrapolation, not 100 measured API requests. Cache hits would reduce the dollar difference: at $1 per million cached tokens, the same token difference is worth about $0.094 per read. Cache writes have their own rate.
Image detail changes the result substantially. Using low is another option when you still need to send images. Applying the documented low-detail image sizing to our assumed PDF page images lowered the estimated PDF reductions to 37.0% for the text report and 66.3% for the scanned packet. The dense standalone image fell to 231 image tokens at low detail—less than its 545-token Markdown version. We did not test whether the model could read that reduced image accurately.
Office files: conversion alone is not the saving
The Office controls tell a different story:
| Test document | Plain-text control tokens | SourceShelf Markdown tokens | Change |
|---|---|---|---|
| Word report, 40 sections | 17,880 | 18,021 | +0.8% |
| PowerPoint deck, 30 slides | 5,370 | 5,674 | +5.7% |
| Excel workbook, 500 data rows | 5,506 | 7,142 | +29.7% |
These controls contain the fixture’s source text; they are not measured API upload counts. They show the cost of headings, provenance, slide boundaries, and Markdown table syntax. Comparing the compressed Office file’s byte size with Markdown would not measure token savings.
For Office documents, SourceShelf’s practical advantage is making the content readable, reusable, and easier to select. In a separate manual selection experiment, keeping four relevant sections from the converted Word report reduced the payload from 18,021 to 1,893 tokens: 89.5% less. Keeping three slides from the converted deck reduced it from 5,674 to 679 tokens: 88.0% less.
Those reductions intentionally exclude material. They are appropriate when the question only needs those sections, and cannot stand in for a whole-document review. The selections were made in the exported Markdown; they were not an automatic relevance feature.
For spreadsheet calculations, keep access to the workbook and an appropriate analysis tool. Markdown tables are useful reference material, but they do not replace formula evaluation or complete spreadsheet analysis. For larger collections, retrieval is another option: OpenAI’s File Search can supply relevant passages instead of putting entire documents into every request.
Review the conversion before relying on it
Fewer tokens are useful only if the information needed for the task survives.
Our checks recovered all expected alphanumeric words, including their repeated occurrences, in the Office files, both standalone images, and the scanned packet. The long text PDF retained 99.65% by that check, with some omissions and reordered text. A word-retention check does not prove that table relationships, reading order, or meaning remain intact.
These clean, synthetic fixtures also do not represent difficult handwriting, complex charts, poor scans, or every business document. SourceShelf’s OCR and format conversion have limitations. Review names, amounts, dates, tables, and reading order; attach the original page or image when it carries evidence the text cannot preserve.
A practical SourceShelf workflow
- Convert locally. Drag your PDFs,
.docx,.pptx,.xlsx, or supported images into SourceShelf on Mac. Text recognition and conversion happen on your device. Local PDF conversion - Inspect the Markdown. Check the information your next task depends on, especially scanned text and tables.
- Build a focused Pack. Include the sources relevant to the question. For narrower tasks, prepare an excerpt from the exported Markdown while preserving useful source labels. SourceShelf reference packs
- Send text and necessary visuals. Supply the Markdown as text to your Astra API workflow. Include individual images when their visual content matters. A Markdown image link alone does not transmit its pixels.
- Reuse the reviewed source. Keep the conversion for future tasks and refresh it when the original changes. Check actual API usage to validate savings in your own workflow.
Local conversion itself does not upload the original document. When you subsequently send Markdown or images to a cloud model, that selected content leaves your device.
Context size, caching, and subscription limits
Large contexts add another reason to choose carefully. Astra prompts above 272,000 input tokens carry higher rates across the full request: twice the input and cache rates, and 1.5 times the output rate. Our test cases stayed below that threshold; reducing a larger request below it could create an additional saving. Count the complete request, including history and tools. Astra model pricing notes
These dollar calculations concern API usage. They do not translate directly into savings on a flat-price ChatGPT subscription or a fixed number of extra Codex tasks.
Prepare once, reuse with care
The useful habit is simple: prepare the source once, check it, and give the model the material the task needs. For text-focused work with PDFs and scans, that can leave substantially more of your token budget available for the work itself.
For another way to keep the result portable, see our guide to Open Knowledge Format. You can keep readable sources and their metadata together as your AI tools change.