Amazon Science’s cover photo
Amazon Science

Amazon Science

Research Services

Seattle, Washington 394,400 followers

The latest news and research from Amazon’s science community. #AmazonScience

About us

Amazon Science gives you insight into the company’s approach to customer-obsessed scientific innovation. Amazon fundamentally believes that scientific innovation is essential to being the most customer-centric company in the world. It’s the company’s ability to have an impact at scale that allows us to attract some of the brightest minds in artificial intelligence and related fields. Our scientists continue to publish, teach, and engage with the academic community, in addition to utilizing our working backwards method to enrich the way we live and work. Follow us on LinkedIn and visit our website to get a deep dive on innovation at Amazon, and explore the many ways you can engage with our scientific community. #AmazonScience

Website
https://www.amazon.science
Industry
Research Services
Company size
10,001+ employees
Headquarters
Seattle, Washington
Founded
2020
Specialties
Artificial Intelligence, Machine Learning, Computer Vision, Cloud, Economics, Sustainability, AI, ML, Conversational AI, Natural Language Processing, NLP, Robotics, Security, Privacy, Information, Knowledge Management, Operations, Scientific Research, Search, Amazon, and Alexa

Updates

  • Verus verifies complex projects with thousands of lines of code and proof in the time prior tools took to verify individual functions. That speed enables an interactive development loop and lets AI agents iterate on proofs faster. Amazon has used this open-source, automated program verifier for Rust to prove correctness of key primitives in the Nitro Isolation Engine and critical internal infrastructure.

  • Eight out of ten LLM judges agree. But how independently did they arrive at that answer? When judges share a prompt template, training lineage, or model family, a majority vote can make evidence look far more convincing than it really is. The vote count inflates confidence without adding independent signal. Amazon researchers introduce dependence-aware label aggregation using Ising models to address this. The method models both individual judge reliability and pairwise dependencies, allowing redundant agreement to be discounted without discarding votes entirely. Across three tasks with 10-judge panels, the approach improved accuracy by 9–14% over weighted-majority-vote baselines.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • Amazon Science reposted this

    View profile for Matt Garman
    Matt Garman Matt Garman is an Influencer

    Honored to be included on this year's TIME100 AI list. This one belongs to our customers and to the teams at Amazon Web Services (AWS). The majority of workloads haven't moved to the cloud yet, and when they do, the opportunity to transform with AI is massive. We're just getting started.

    View organization page for AWS Newsroom

    29,245 followers

    AWS CEO Matt Garman has been named to the TIME100 AI list for 2026. 🎊 Matt joined as an MBA intern in 2005, before AWS even had a name, and twenty years later he leads the world's largest cloud provider, now on a $169 billion annualized revenue run rate. Congratulations, Matt. Read the full TIME100 AI profile: https://lnkd.in/eSA6V-rw

    • No alternative text description for this image
  • How did a model upgrade make agents worse? SOP-Bench is a new benchmark that pairs genuine enterprise procedures with functioning tools and ground-truth answers across 12 industries and 2000+ tasks. With a reasoning-style (ReAct) agent, the newer Claude 4.5 family scored lower than the older Claude 4 family. A routine upgrade can lower an agent's success rate with no obvious signal that anything changed. The only reliable way to catch it is to test on the procedures the team actually runs. With SOP-Bench, agents earn their scores by completing the work rather than producing text an automatic grader happens to like.

    • No alternative text description for this image
  • Ten years ago, Amazon's Automated Reasoning Group set out to use mathematical logic to prove AWS systems work correctly. Today, their production services process billions of queries daily, powering tools millions of customers rely on, from IAM Access Analyzer to Amazon Inspector to Bedrock Guardrails. One project stands out: they proved correct and seamlessly replaced the AWS authorization engine, which handles one billion API calls per second, and verified it against quadrillions of production authorizations. Now the same formal-verification techniques that secured cloud infrastructure are being applied to AI, setting boundaries for autonomous agents and validating AI-generated content with up to 99% verification accuracy.

  • AWS Trainium Frontier challenges researchers to train language models from scratch on purpose-built AI chips. The NeurIPS 2026 competition is open for registration, limited to 100 teams. Teams optimize across the full stack, including model architecture, optimizer, training loop, and custom hardware kernels. No prior Trainium experience is required. Prizes include $25K for first place, co-publication with Annapurna Labs researchers, and a chance to present at an event during NeurIPS 2026 in Sydney. Register by September 30: https://amzn.to/3TL03EA

  • View organization page for Amazon Science

    394,400 followers

    Announcing the 34 recipients of the Amazon Research Awards Build on Trainium program, a $110 million credit initiative supporting AI research at 30 universities. This cycle focused on Responsible AI, inviting proposals in AI safety and alignment, multi-lingual language models, representation engineering, sustainability and small language models, and deep learning models for synthetic data generation. Awardees have access to more than 700 Amazon public datasets, AI/ML services and tools, and AWS Trainium resources including tutorials and hands-on sessions.

  • AI's full potential is bottlenecked by an efficiency problem that spans the entire stack. At Berkeley RDI's Agentic AI Summit, Amazon SVP Peter DeSantis traced how the predictable memory and compute flow of AI models led Amazon to build Trainium on a systolic array architecture, stripping out flexibility those workloads don't need while preserving what matters. But no single chip will power the next decade. As AI workloads keep evolving, Peter sees a growing and more diverse hardware ecosystem ahead. Annapurna Labs' John Liu followed with a deep dive on using agentic looping to optimize models on Trainium. Key takeaways: test agents on data they've never seen, verify visibility scope before fixing multi-agent failures, and recognize when existing rules are causing the very failures you're trying to patch.

  • Multitask training objectives fight each other at every gradient step. What if you stopped forcing them to share? ControlG reframes multitask coordination as a temporal allocation problem, dedicating computational capacity to one objective at a time instead of blending conflicting gradients per step. A proportional-integral-derivative (PID) controller decides which objective needs attention next, operating across three time scales: estimating per-objective difficulty, optimizing per-epoch allocation via log-hypervolume sensitivity, and tracking the plan with feedback loops. Amazon researchers applied ControlG to the problem of graph self-supervised learning. Across nine graph datasets, ControlG achieves average ranks of 1.4, 1.9, and 1.8 for node classification, link prediction, and node clustering, exceeding all baselines.

Affiliated pages

Similar pages