Ema
Software Engineer, Full-Stack
- Location
- San Francisco Bay Area, CA, US
- Employment type
- Full-time
- Posted
- 1 October 2026 (today)
Checked todayApplications go to the employer, never to RoleSprint
About this role
ABOUT EMA
Ema builds AI Employees for HR, IT and Finance. Our AI Employees take on the busy work across the employee experience, from recruiting, onboarding and benefits to IT support, invoice processing and payroll, so people can spend their time on work that needs them. Founded by former executives from Google, Coinbase, Flipkart, and Okta, our team includes engineers from premier tech companies and graduates of Stanford, MIT, UC Berkeley, CMU, and IITs.
We are backed by industry leading investors including Accel, Naspers/Prosus, Creaegis, Section32, and angels like Sheryl Sandberg and Dustin Moskovitz. Headquartered in Silicon Valley and with offices in London, Bangalore and Vancouver, Ema is at the frontier of what Agentic AI can do in production. We ship real systems that run real business processes at scale.
BUILD AGENTS THAT LEARN FROM THE WORK THEY DO.
Ema builds AI employees that carry out complex workflows across enterprise applications. Our ML team works on the loop that makes them better: production traces become data, data becomes training and evaluation, and better agents produce better traces. The hard part is deciding which intervention will improve behavior in the next real workflow.
THE PROBLEM SPACE
- Harnesses and inference-time compute. Design context, tools, skills and orchestration for multi-step agents, including work across documents, slides, images, audio and video. Test where extra reasoning, search or verification earns its latency and cost. Build self-improvement loops with explicit permissions, evaluation gates and rollback.
- Agent post-training. Curate trajectories for SFT, optimize preferences, or run RL on real agent tasks. Investigate methods such as DPO, GRPO or DAPO where they fit; compare process and outcome supervision, shape rewards, and distill useful frontier behavior into smaller models. Measure whether gains transfer beyond the training environment.
- Environments and rewards. Turn enterprise workflows into reproducible training and evaluation environments: fixture tenants, simulated users who may get impatient and leave, and rewards grounded in verifiable outcomes. Find the shortcuts an agent can exploit before a training run optimizes for them.
- Data engines and evaluation. Mine production agent-steps for failures; build curated corpora and useful synthetic augmentation. Calibrate judges against human labels, construct behavior-level benchmarks from real workflows, and quantify data quality, performance uplift and reliability across stochastic runs.
- Retrieval, memory and context graphs. Connect enterprise information with user- and tenant-level learnings. Separate failures of retrieval from failures to use retrieved context; test what to retain, update and retrieve so that past experience improves the next decision.
- Quality per dollar. Build and evaluate routing, ensembles, caching and small-model specialization. Measure downstream task success alongside latency and cost; a cheaper model is useful only if the complete agent still succeeds.
You’ll go deep in a subset of these areas. Projects combine applied research with the engineering needed to make the result work in production.
OWN THE EXPERIMENT AND THE SYSTEM
Most projects develop over roughly four to six months, with useful improvements shipping along the way. You’ll define the problem and baseline, build the data or environment needed to test it, run experiments, and own serving and integration. Follow the system through deployment, monitoring and failure analysis until it is ready for a clear engineering handoff.
You’ll work with researchers and engineers across the Bay Area, Vancouver and India. We’re hiring from junior through senior levels, with project scope matched to your experience.
WHAT WE’RE LOOKING FOR
- Depth you can defend. Substantial work in at least one of agent/tool-use systems, post-training, reward modeling or RL environments, retrieval and memory, or evaluation design. Be ready to explain the mechanism, the alternatives you rejected and the failure modes you found. One area you can teach us beats five you’ve touched.
- Evidence of zero-to-one ownership. A system, model or research project you took from an ambiguous problem to a working result. We want to understand your contribution, the tradeoffs you made, and what changed when the work met real users or realistic tasks.
- Statistical judgment. You can size an experiment, choose meaningful baselines and held-out tests, and account for variation across tasks, seeds and repeated runs. You can distinguish a real improvement from judge bias, data leakage or a benchmark shortcut.
- Production engineering judgment. You can debug across the model and system boundary, isolate a failure, and turn the result into maintainable production code.
- Honest measurement. You would rather retire your own approach after a clean negative result than ship an improvement that disappears under a stronger evaluation.
A master’s or PhD in a relevant field, or equivalent work or research experience. Papers, substantial open-source contributions, trained models and well-documented experiments can demonstrate that depth.
EXPERIENCE THAT WOULD BE ESPECIALLY USEFUL
- Open-model post-training with TRL, veRL, OpenRLHF, or similar or a custom loop, especially debugging reward hacking or unstable optimization.
- Interactive agent environments/harnesses for software engineering, web or tool use; large-scale trace analysis, data curation or synthetic generation.
- Designing systems to support complex, long-horizon agent work across a multitude of modalities and platforms
- Serving with vLLM or SGLang, distillation, quantization, or multi-node GPU training.
- Practical security work on prompt injection, data governance or permission boundaries for agents that can act and improve themselves.
These are project-specific strengths; post-training experience is optional, and no candidate needs the entire list. Bring a repository, paper, model or technical write-up that lets us examine how you think and what you built.
Compensation offered will be determined by factors such as location, level, job-related knowledge, skills, and experience. Certain roles may be eligible for variable compensation, equity, and benefits.
Ema Unlimited is an equal opportunity employer and is committed to providing equal employment opportunities to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, sexual orientation, gender identity, or genetics.
Work location
- San Francisco Bay Area, CA, US
Related jobs
Ready to make a decision?
This role is either worth your time or it isn’t.
Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.
Nothing is submitted automatically. You choose what happens next.
About this listing
Advertised by Ema and published on Ashby, the applicant tracking system they use.
RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.
Published 1 October 2026, last checked today. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.