Physical Superintelligence
Member of Technical Staff, ML Engineer
- Location
- Boston, MA, US
- Arrangement
- Hybrid
- Employment type
- Full-time
- Level
- Staff
- Posted
- 1 August 2026 (about 2 months ago)
Checked 22 days agoApplications go to the employer, never to RoleSprint
About this role
OVERVIEW
Physical Superintelligence is a startup with roots at Google, NVIDIA, Harvard, Meta, MIT, Oxford, Johns Hopkins, Cambridge, and the Perimeter Institute building AI systems to discover new physics at scale.
Our mission is to discover and commercialize transformative physics breakthroughs at scale with artificial superintelligence, safely, verifiably, and for broad public benefit.
The last century's golden age of physics gave us transistors, lasers, and nuclear energy. We believe artificial superintelligence will unlock the next one. We're creating the infrastructure to industrialize scientific discovery and usher in this new era.
We have one product: new physics, at scale.
ROLE
We are seeking a Member of Technical Staff, ML Engineer to build and run the training and inference systems that turn Core AI's research into things that work at scale, and that the rest of Engineering can build on.
RESPONSIBILITIES
- Own the training and inference infrastructure that Core AI depends on: distributed training jobs, GPU scheduling, and model-serving systems (vLLM, SGLang, or comparable) for both proprietary models and self-hosted inference.
- Build the tools and abstractions AI researchers use to launch training runs, iterate on inference providers, and route workloads across models, so a researcher's time goes into the science instead of the plumbing.
- Partner with Engineering on the shared platform: capacity planning, observability, and reliability for GPU and inference infrastructure, so training and serving hold up to the same production bar as everything else we ship.
- Debug and harden the training and inference stack under real load. Egress failures, stalled retries, and routing edge cases are your problem to close, not someone else's ticket.
- Stay hands-on. You write the code, not just the design doc, and you are the first call when a training job stalls or an inference path breaks.
WHAT WE'RE LOOKING FOR
- Three or more years building and operating ML training or inference infrastructure in production, at a company that trains or serves models at meaningful scale.
- Hands-on experience with distributed training (multi-GPU or multi-node, using PyTorch, Ray, or comparable) and model-serving systems (vLLM, SGLang, Triton, or comparable).
- Strong software engineering fundamentals. You can build a service that other engineers and researchers depend on every day, not a script that worked once.
- Enough ML fluency to work productively with AI researchers: you understand training loops, reward signals, and inference-time behavior well enough to debug them, even without designing the algorithms yourself.
NICE TO HAVE
- Experience building internal platform tools such as training-as-a-service APIs, inference gateways, or job schedulers.
- Background in GPU infrastructure, CUDA, or performance engineering for ML workloads.
- Experience with cloud infrastructure (GCP, AWS) and infrastructure as code (Terraform or comparable).
- Prior work embedded alongside a research team, turning research code into production systems.
HOW WE WORK
We hold a high technical bar and give people full ownership of their work, from spec to ship to on-call. We write contracts before logic, test against real systems instead of mocks, and favor simple designs that ship over clever ones that do not. Our development process is AI-native: we work with agentic coding tools daily, write specs that are legible to humans and agents alike, and lead with leverage.
LOCATION AND COMPENSATION
This role is based in Boston. We will consider remote candidates on a case-by-case basis. We offer competitive compensation including salary, benefits, and meaningful early-stage equity. We evaluate on technical breadth, systems thinking, ML infrastructure depth, and shipping velocity. We are an equal opportunity employer and value diverse perspectives in building platforms for AI-driven discovery.
Work location
- Boston, MA, US
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Published on Ashby under the board identifier Psi, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 1 August 2026, last checked 22 days ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.