Ema
AI Resident
- Location
- San Francisco Bay Area, CA, US
- Arrangement
- Hybrid
- Employment type
- Full-time
- Posted
- 31 July 2026 (about 2 months ago)
Checked todayApplications go to the employer, never to RoleSprint
About this role
ABOUT EMA
Ema builds AI Employees for HR, IT and Finance. Our AI Employees take on the busy work across the employee experience, from recruiting, onboarding and benefits to IT support, invoice processing and payroll, so people can spend their time on work that needs them. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs.
We are backed by industry leading investors including Accel, Naspers/Prosus, Creaegis, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production. We ship real systems that run real business processes at scale.
The residency
You own one hard problem end to end. You write the proposal, build the system, design the evaluation, ship behind a gate, and finish with a write-up of what turned out to be true, including the parts that didn't work. You'll sit in the production codebase with a senior mentor and real production data. Recent residents have shipped self-improving harnesses, inference-cost work, agent memory, and eval infrastructure. Your project gets scoped with you, not handed to you.
The problem space
The loop we care about: production traces become data, data becomes training and evaluation, and better agents produce better traces. Projects live somewhere on that loop.
- Harness and inference-time work. Context engineering, tool and skill design, orchestration, and deciding where extra inference compute actually pays. Self-improvement loops run behind hard fences.
- Post-training for agents. SFT on curated trajectories, preference optimization, RL on real agent tasks. Reward design where outcomes are verifiable, process vs. outcome supervision, distilling frontier behavior into cheaper models.
- Environments and rewards. Turning enterprise workflows into training and eval environments: fixture tenants, simulated users (some of whom get impatient and leave), verifiable rewards, and defenses against reward hacking. Agents will exploit a lazy grader.
- Data engines. Mining production agent-steps into training and eval corpora: failure mining, labeling with calibrated judges, synthetic augmentation that stays useful.
- Evaluation. Behavior-level benchmarks from real workflows, LLM judges calibrated against human labels, reliability statistics for stochastic agents.
- Efficiency. Routing, ensembles, caching, small-model specialization. Quality per dollar is a research metric here.
What we're looking for
- No specific degree required. Strong undergrads, grad students, and self-taught builders are all welcome; what matters is demonstrated depth in ML or agent systems.
- Solid ML fundamentals and strong engineering: Python, PyTorch, and the discipline to ship in a large production codebase.
- Real depth in at least one of: post-training (SFT/DPO/GRPO-family RL), reward modeling or LLM judges, agent and tool-use systems, retrieval and memory, eval design. One area you can teach us beats five you've touched.
- Statistical literacy. You can size an experiment, and you know 25 samples at one seed is a datapoint, not a result.
- Honest measurement as a habit. You'd rather kill your own feature with a clean experiment than ship it on a hunch.
Nice to have
- Hands-on post-training with open models (TRL, veRL, OpenRLHF, or your own loop). Bonus points if you've debugged a reward-hacked run.
- Built or trained in interactive agent environments (SWE, web, or tool-use gyms).
- Large-scale trace analysis, data curation, or synthetic data work.
- Serving and efficiency experience: vLLM/SGLang, distillation, quantization.
- Multi-node GPU training, or the infra fluency to get there fast.
- Publications, open-source work, or writing that shows how you think.
- Security instincts: prompt injection, data governance, why a self-improving agent needs a fence.
Logistics
- SF Bay Area, on-site/hybrid, half/full-time for the term. Flexible start.
- Salary: $4,000 month
Compensation offered will be determined by factors such as location, level, job-related knowledge, skills, and experience. Certain roles may be eligible for variable compensation, equity, and benefits.
Ema Unlimited is an equal opportunity employer and is committed to providing equal employment opportunities to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or genetics.
Work location
- San Francisco Bay Area, CA, US
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Advertised by Ema and published on Ashby, the applicant tracking system they use.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 31 July 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.