Log inSign up
Anthropic
1,684 posts
Anthropic profile banner
@AnthropicAI

Anthropic

@AnthropicAI
We're an AI safety and research company that builds reliable, interpretable, and steerable AI systems. Talk to our AI assistant @claudeai on claude.ai.
anthropic.com
Joined January 2021
2
Following
1.6M
Followers
RepliesRepliesRepostsRepostsMediaMedia

Log in or sign up for X

See what’s happening and join the conversation

Continue with phone
or
Log in with username or email
Terms·Privacy·Cookies·Accessibility·Ads Info·© 2026 X Corp.
  • @AnthropicAI
    Anthropic
    @AnthropicAI
    9h
    New research: Training a Misaligned Reward Seeker What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an
    144
  • @AnthropicAI
    Anthropic
    @AnthropicAI
    11h
    We’re sharing an update on our alignment and security efforts. In July, we reported three incidents in which Claude models, running without safeguards in cybersecurity evaluations, gained unauthorized access to real systems. In a new post, we describe: 1. How we’ve secured
    Hand with geometric shapes constructing a complex abstract form on white background
    Improving our alignment and security practices
    From anthropic.com
    333
  • @AnthropicAI
    Anthropic
    @AnthropicAI
    Aug 28
    New Fellows Research: Can Claude autonomously align other AIs? We gave Claude 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own. It worked surprisingly well.
    Geometric staircase steps ascending vertically with incremental progression
    Automated researchers can reliably mitigate alignment failures
    From anthropic.com
    364
  • @AnthropicAI
    Anthropic
    @AnthropicAI
    Aug 27
    Today, we're kicking off the first phase of the research preview for Model Hardware Standard (MHS): a new standard for AI agents to safely operate physical equipment in scientific research and advanced manufacturing. Read more: anthropic.com/news/model-har…
    00:00
    506
  • @AnthropicAI
    Anthropic
    @AnthropicAI
    Aug 26
    For the first time, we’ve given external researchers a way to study AI’s impacts using real, privacy-preserved Claude usage data. To date, this work has only been possible within AI labs. We can’t tell the whole story alone, so we opened up our tools.
    Desk lamp beside stacked books on wooden surface with academic study materials
    Enabling independent research on how people use Claude
    From anthropic.com
    202