5buyai × AIFUNS
Promoted
← Back to feed
Anthropic
@AnthropicAI
We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on https://t.co/FhDI3KQh0n.
1.6MFollowers
2Following
20Cached
2Videos
Amazon Web Services Amazon Web Services Lightreel AI Lightreel AI ElevenLabs Developers ElevenLabs Developers Figma Figma Netflix Netflix DigitalOcean DigitalOcean Notion Notion Runway Runway Spline Spline Microsoft Azure Microsoft Azure YouTube YouTube Telegram Messenger Telegram Messenger Google Ads Google Ads PixVerse PixVerse LovartAI LovartAI OpenClaw🦞 OpenClaw🦞 Canva Canva MiniMax (official) MiniMax (official) GitHub GitHub Spotify Spotify Red Hat Red Hat MiniMax Design (H3) MiniMax Design (H3) X X Google Cloud Google Cloud Instagram Instagram Pika Pika Cursor Cursor OpenAI OpenAI Twitch Twitch Midjourney Midjourney
All 🖼 Media 🎬 Video
Could a model one day align its stronger successors? As a first test, we had Sonnet 5 post-train an early checkpoint of Opus 4.8, a more capable model. It reached safety scores approaching those of production Opus 4.8, which went through our full alignment training.
Across 10 alignment failures, Claude reliably improved safety scores without degrading capabilities. Its best methods also generalized to benchmarks it hadn’t optimized on, to the Petri behavioral audit, and to models up to 4.7x larger.
Claude “hill-climbed” safety benchmarks for common misalignments like deception or sycophancy, with one constraint: it had to preserve general capabilities. We then tested its best methods on held-out benchmarks to see if they'd generalize.
MHS currently best covers lab and manufacturing equipment. Many developers are already using Claude Code to operate hardware like boards and cameras; our research preview will help us extend MHS to these devices, so they can all work under one interface.
There’s more to learn before we open source MHS. LLMs still lack physical intuition, having learned about the physical world from text and images. The research preview will let us build more safety evaluations and strengthen protections for using AI in the physical world.
In early testing, AI agents used MHS to: Run a drug-discovery experiment with real-time error handling at Genentech Compress an imaging experiment from weeks to a day at HHMI Janelia Research Campus Improve laser stabilization on QuEra's quantum computers from 58% to 99.3%
Connecting AI to hardware requires days or weeks of bespoke integration, with no standard way for agents to operate equipment safely. MHS cuts integration to hours or minutes, provides an interface that makes devices discoverable, and enables agents to operate them safely.
Now, we want to scale this research model. If you're a researcher and would like access to our tools to pursue work you can't otherwise do today, we’d like to hear from you. You can express interest here: https://forms.gle/rmLjTvibven9CmDFA
The other two studies are ongoing: HIP Lab is studying how Claude's behavior relates to how people feel when using AI, while METR is estimating real-world productivity gains from coding agents. We'll share more from both soon.
Three research groups—Stanford’s Social and Language Technologies lab, Oxford’s Human Information Processing Lab, and METR—designed independent studies to analyze the aggregated outputs from 250,000 https://Claude.ai or Claude Code conversations between April and May 2026.
One of our highest priorities remains launching an access program for scientists to use our most capable models. We expect to share more on this soon. Opus 5 remains our most capable model available for life science research.
Aggregated from the public X timeline; copyright belongs to the original author. This site is not affiliated with this account.
💬 Need a hand?
Purchase / payment / account help — chat with us →

Announcements

If you have a credit card, you can register an account on this site, use your credit card to top up your balance in the personal center, and then use the balance to pay for purchasing our products or services.