Baseten’s cover photo
Baseten

Baseten

Software Development

San Francisco, CA 41,481 followers

Own your inference.

About us

Inference is everything. Baseten is an AI infrastructure platform giving you the tooling, expertise, and hardware needed to bring great AI products to market - fast. Our proprietary Inference Stack utilizes the cutting-edge of performance research combined with highly performant and reliable infrastructure to give you out-of-the-box global availability with 99.99% of uptime.

Website
https://www.baseten.co/
Industry
Software Development
Company size
201-500 employees
Headquarters
San Francisco, CA
Type
Privately Held
Specialties
developer tools, software engineering, artificial intelligence, and machine learning

Products

Employees at Baseten

View 413 employees at Baseten

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

See all employees

Locations

Updates

  • Baseten reposted this

    the cheapest way to get Opus 5-level agentic performance right now is an open model. GLM-5.3 is 743B params with 40B active. that's about a quarter the size of Kimi K3, and it matches or beats it on intelligence. the benchmarks are clear. the harder questions are whether it holds up on your tasks and what a task actually costs you at production volume. that's what we're covering wednesday 9/2 at 11am PT. members of our technical staff on serving GLM-5.3, benchmarking it on your own evals, and working out per-task cost for your workloads. register here: https://lnkd.in/ee6tWqNt

    • No alternative text description for this image
  • Baseten reposted this

    As my internship at Baseten wraps up I wanted to show what I've been working on! Over the past year, agents have shown to be remarkably capable of generating high-performance kernels from scratch. But generating an isolated kernel is very different from improving and shipping an optimization for production within a serving engine. We built a framework to bridge that gap, offering a glimpse into the future of kernel engineering. https://lnkd.in/gR89EfWK

    • No alternative text description for this image
  • Baseten reposted this

    Last night Baseten hosted Bryan Johnson and Sarah Guo in San Francisco. Our thesis is that obsessives move the world forward, and few people are more obsessive than Bryan. So I assumed the night would be about intensity: how to sustain it, how to get more out of yourself. He went the other direction entirely. Bryan came to this from the other side: He sold his business, went through a divorce, and was chronically depressed before any of this started. His message now is that we've been trading our health for money, success, and power without ever pricing the trade. His line: you're not taking care of yourself, you're not writing good code. One my favorite points was about how our thinking narrows when we don't have alternatives: how we accept chronic illness, death because there is no other choice. He compared it to body positivity - a message that has largely faded since GLP-1's arrived. I left inspired to take better care of myself, and to stop trading my health so easily. Thank you to Bryan and Sarah for a conversation none of us expected. And thank you Navya Gudimetla, Lauren Lechuga and the Immortals team!

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • We're proud to be the fastest inference provider on Artificial Analysis, OpenRouter, and Hugging Face for GLM-5.3-Flash, at 122+ TPS. All served from the US only, starting on day 0, with ZDR by default. Stay tuned for updates as our engineers continue to optimize GLM-5.3-Flash for even higher throughput and lower latency.

  • Baseten reposted this

    Mika Okamoto and I are releasing PACT (Pressure-Applied Compliance Testing): a benchmark for whether enterprise AI assistants keep following workplace rules when breaking them is convenient. We tested 23 models (18 open-weight, 4 closed) on 3,364 multi-turn trials across regulated domains like HIPAA, hiring law, and GDPR. Each trial gives the assistant a standing rule, makes the violating option the easy one, and adds ordinary corporate pressure: a deadline, a manager's verbal OK, "my colleague did it and nothing happened." PACT spun out of our #AIES2026 paper on why AI agents break rules; this time we turned the question into a full benchmark. The two highest-scoring models were open-weight. Kimi K2.7 (0.944) and Qwen3.6 27B (0.943) finished statistically tied, ahead of every closed model we tested. Bigger isn't better here; a 27B you can run yourself co-leads the board, and open models are competitive on all six dimensions (baseline, pressure resistance, transparency, and others) we measure. Not one model aced it. One sentence of pressure raised violation rates 65%, and the best model still missed 1 decision in 18. If you're deploying in a regulated workflow, evaluate on your own workload, and don't assume the closed-source frontier is the frontier for rule-following. Leaderboard, interactive trials, dataset, paper, code: https://lnkd.in/gPmB73Ns Baseten funded the inference for all ~232,000 trials. An inference company paying to measure how models behave, not just how fast they serve, is why I love working here.

    • No alternative text description for this image

Similar pages

Browse jobs