AI

Introducing Claude Fable 5.1 and GPT-6 Astra

In the first week of September, two new flagship AI models were released within days of each other, each arriving with bold claims about cost, capability and intelligence. Anthropic released Claude Fable 5.1, and OpenAI released GPT-6 Astra. Between them, they power a lot of the content, coding, and research work happening online right now.

Anthropic launched Claude Fable 5.1 on 1 September. It’s the company’s most capable publicly available model, designed for long, complex jobs that run for hours, days, or even weeks. Two days later, OpenAI released GPT-6 Astra as a limited preview, opening it up to paid users the following day. The two launches lean on very different selling points:

  • Fable 5.1 is being sold on cost. Anthropic has cut the price of cache reads by 75%, from $1.00 to $0.25 per million tokens. Tokens are the small chunks of text a model reads and writes, and they’re how usage is billed. Caching lets the model reuse information it has already processed, such as your instructions or reference documents, instead of paying full price to read them again.
  • GPT-6 Astra is being sold on capability. It can operate a computer much like a person would, browsing websites through an in-app browser or Chrome extension and working inside desktop apps on macOS and Windows.

 

Claude Fable 5.1 Cuts the Cost of Cached Content

Anthropic says the cheaper cache reads bring costs down by around 25% on typical workflows and up to 45% on agentic ones, where the model works through long, multi-step tasks on its own. Those savings are real, but they depend on how you use the model. Per-token pricing hasn’t changed – if your prompt requires information that isn’t already cached, you’ll pay the old rate to pull in that extra information.

The saving comes with conditions… Most new models let you choose how hard they think before responding. Fable 5.1 has five effort levels: low, medium, high, xhigh, and max. The effort setting moves the bill more than model choice. Artificial Analysis measured an 11-fold gap in token use between the lowest and highest effort settings on the same model – 13.1 million output tokens at low effort vs. 143.7 million at max. Push Fable 5.1 to maximum effort and it actually costs more per benchmark task than the version it replaced, at $3.76 against $3.14.

 

Bigger is not the same as better suited

The more useful question is not which model is strongest, but which one fits the job. Anthropic doesn’t recommend its own flagship as a default. Its guidance is to start on Sonnet 5.5 and move up to Fable 5.1 only for long-horizon work, or when a cheaper tier has already been tested and fallen short. Matching the tier to the task is a sensible starting point:

Model tier Best suited to Examples
Haiku High-volume, repetitive work Tagging products, sorting enquiries, first-draft meta descriptions
Sonnet Everyday tasks Drafting emails, summarising reports, routine content edits, agentic coding
Opus Heavier judgement calls Strategy documents, detailed analysis, complex editing
Fable Sessions that run for hours Large research projects, site migrations, multi-stage builds

Point a top-tier model at a small rewrite and it will do the job and burn through tokens doing it. It has no way of knowing the task was small. Anthropic’s documentation makes the point directly: “tuning effort is often a better lever than switching models.”

 

OpenAI’s GPT-6 Astra

OpenAI President Greg Brockman described the release of GPT-6 Astra as a “generational leap” that signals the arrival of AGI (Artificial General Intelligence) – a turning point where AI is broadly as capable as humans at complex tasks. Astra’s standout feature is computer use. Other models need a custom integration for every piece of software they touch. Astra instead works across browsers, spreadsheets, websites, and desktop apps much as a person would, filling in forms, moving between web pages and producing finished documents and presentations.

OpenAI’s examples range from building video game scenes to ordering food and running job searches. In its launch demo, staff directed the model entirely by voice, turning a simple drawing into a 3D game and creating an eBay listing in minutes. OpenAI also describes Astra as its strongest model yet for coding, maths and science. It scored 66% on ARC-AGI-3, a test of how well a model reasons through problems it hasn’t seen before, using the benchmark’s standard setup.

On OSWorld V2-Offline, a benchmark for desktop application work, Astra scored 72.6% against 65.7% for its predecessor, and cut average task completion time from roughly 75 minutes to 40. For marketing teams, that could eventually mean handing over jobs like updating product listings or pulling performance reports from several platforms at once. Astra is genuinely impressive and we’re excited to see what computer use makes possible. Still, we’d urge a little caution. Its benchmark figures rely partly on OpenAI’s own testing setup, so real-world results may vary.

 

Cybersecurity and Recurrent Depth Reasoning

The bigger concern around GPT-6 Astra is transparency. Earlier AI models “showed their work”, writing out each step before arriving at an answer. Newer approaches do more of that thinking internally, a bit like doing mental maths instead of working a sum out on paper. Recurrent depth reasoning is faster and uses fewer resources, but, when you can’t see how a model reached its answer, you can’t check whether its logic held up. Oleksandr Yaremchuk of Manifold Security, speaking to TechRadar, found Astra concealed its reasoning in the majority of cases tested:

“OpenAI is calling Astra its most aligned model yet, even as its chief scientist admits monitorability is getting harder as models get more capable. […] That matters because Astra isn’t staying inside OpenAI’s test environment. It’s going to run as an agent on employee laptops and in the browser, holding real credentials, inside companies that have no way to watch what it does once it’s there. A model that explains itself less isn’t more aligned, it’s just harder to catch when it goes wrong.”

A specialist can usually spot a wrong answer in their own field. Someone without that knowledge has nothing to cross-check against. You can ask the model to explain itself afterwards, though that gives you an after-the-fact account, which may not reflect what happened along the way. Enterprise workspaces can restrict which sites and applications GPT-6 Astra touches, and require approval before it acts. Those controls are definitely worth setting up before the capability is switched on, not after.

herdl-search-engine-optimisation-pay-per-click

 

The Last Word

AI is moving quickly, and it’s easy to feel pressure to chase every release, but you don’t need to. We advise to let the dust settle after a new launch and early adopters have had time to find the flaws in a model before rushing to trust it. The businesses getting the most from these tools are the ones using them thoughtfully. And the newest, most powerful model isn’t automatically the right choice. If you’re building AI into your marketing or day-to-day operations, here’s where we’d start:

  • Begin on the cheapest tier that does the job well and step up only when it falls short
  • Keep effort settings low for routine work and save higher settings for complex tasks
  • Keep reusable instructions and reference material consistent so caching can do its job
  • Check facts, figures and quotes before anything goes live, especially when an AI model hides its reasoning
  • Keep someone with real subject knowledge involved in reviewing the output.

If you’d like to talk through where AI fits into your marketing, we’d love to have that conversation. Get in touch.

Want more?

Ready to grow your business?

Whether you’re looking to improve your website, boost your visibility, or scale your digital marketing, we’re here to help. Get in touch using the form below or call us on 0116 3400 442

"*" indicates required fields

This field is for validation purposes and should be left unchanged.