GPT-6 Astra vs Claude Opus 5.5: Filling a 1M Context
September 2026 gave us two new flagships three weeks apart. OpenAI shipped GPT-6 Astra on September 3–4, then added the lighter Sol and Luna on September 22. The same day, Anthropic released Claude Opus 5.5.
Most comparisons stop at the benchmark table. This one asks a more practical question, the one you hit the moment you actually use a million-token window: what do you put in it, and what does it cost to fill?
Because a bigger window doesn’t make your sources cleaner. It just lets you pay for more noise.
The two windows, side by side
Here is what matters for long-context work, from the published specs and list prices at the time of writing:
| Claude Opus 5.5 | GPT-6 Astra | |
|---|---|---|
| Context window | 1,000,000 tokens | 1,050,000 tokens |
| Max input in one prompt | 1,000,000 | ~922,000 (the rest is reserved for output) |
| Input price (per 1M tokens) | $4 | $10 |
| Output price (per 1M tokens) | $20 | $50 |
| Long-prompt surcharge | None, flat rate | Whole request billed ~2× input, 1.5× output above 272K input tokens |
Two things jump out.
First, the “bigger” window on paper is not the bigger prompt in practice. Astra’s window is 5% larger, but because part of it is held for output, Opus 5.5 accepts the larger single input.
Second, and more important for your wallet: Astra has a price step at 272K input tokens. Cross it and the entire request, not just the overflow, gets repriced. A prompt of 280K tokens costs roughly twice what a prompt of 270K does.
That single number changes how you should prepare context for it.
Where the tokens actually go
Say you want a model to reason across 40 web pages: competitor docs, a few research articles, some forum threads. The naive route is to paste or fetch the raw pages.
Raw HTML is mostly not content. It is navigation, cookie banners, scripts, inline styles, tracking attributes, and footer links repeated on every page. A 5,000-word article can easily weigh 30,000–50,000 tokens as HTML. The same article as clean Markdown is typically 80–90% smaller, often under 7,000 tokens.
Run that across 40 pages:
- Raw HTML: ~1.4M tokens. It doesn’t fit either window. You start cutting pages.
- Clean Markdown: ~250K tokens. It fits both, and it stays under Astra’s 272K step.
Same sources, same question. One version doesn’t run; the other runs on either model, at base price.
Noise doesn’t just cost money, it costs answers
Long-context models have become very good at finding a needle in a haystack. They are still better when there is less hay.
When half the prompt is menus and boilerplate, three things degrade:
- Attribution. Repeated nav text and “related articles” blocks create false matches. The model quotes a sidebar instead of the article.
- Structure. HTML hides hierarchy in
divsoup. Markdown keeps it explicit:#headings, lists, tables and code blocks the model reads natively. - Cache reuse. Both vendors discount cached input heavily. Clean, stable Markdown files make a much better cache prefix than pages that change their ads on every fetch.
A simple routine for 1M-token work
This is the workflow we use ourselves, and it works the same with either model:
- Collect once, as Markdown. When you read something worth keeping, save it as a
.mdfile instead of a bookmark. With Minibase that’s one click in Chrome: the page becomes clean Markdown with title, source URL and date in the frontmatter. - Group by project. One folder per question you’re working on. Forty files in a folder is a context pack.
- Count before you send. Markdown makes this honest: file size divided by roughly four gives you a token estimate. If you’re aiming at Astra, keep the pack under 272K.
- Feed the pack, not the internet. Point the model at the folder. With Minibase Vault connected to Claude, Claude reads your saved files directly through MCP, with no copy-paste.
- Keep the pack for next time. Next month, when the next model ships, the same folder works unchanged.
So which one should you use?
For long-context research and document work, the published numbers make Opus 5.5 the easier default: a flat price across the whole window, the larger single prompt, and a lower rate per token. Astra is a strong model with genuine wins, especially in computer use, and it’s worth using if you’re already on OpenAI’s stack. Just keep your prompts under the step.
But the honest takeaway is that the choice of model matters less than the choice of input. Both windows reward the same discipline: fewer, cleaner tokens. The team with 250K tokens of well-structured Markdown will get a better answer, faster and cheaper, than the team pasting 900K tokens of HTML into whichever model topped this week’s leaderboard.
Models change every few months. Your sources don’t have to.
Related reading
Continue reading
The URL-to-Markdown API built for AI agents and RAG
LLMs read Markdown, not HTML. How to feed agents and RAG pipelines clean web content with one API call, and why a tiered engine beats scraping.
Why Markdown Is the Best Format for LLMs and AI Agents
Markdown reduces token usage by up to 10x compared to HTML. Learn why AI agents and LLMs prefer Markdown for context and how to optimize your AI workflows.
Graph Engineering, Explained: Loops to Knowledge Graphs
Graph engineering started as a joke in July 2026 and stuck. Here are its three meanings, and why your Markdown vault is already most of the graph.
Jack Dorsey's Buzz: Agents, Git, and Markdown
Block's Buzz puts AI agents in team channels with signed identities and Git built in. What it gets right, and why Markdown still owns the memory layer.