Skip to content

LlamaIndex

Member of Technical Staff, Applied Research

Location
San Francisco, CA, US
Arrangement
Hybrid
Employment type
Full-time
Level
Staff
Posted
8 July 2026 (about 2 months ago)

Checked about a month agoApplications go to the employer, never to RoleSprint

About this role

Join us and help shape the future of AI by defining the narrative around document understanding.

ABOUT THE ROLE

We are looking for an AI Research Engineer to join our document understanding team.

This role is ideal for someone who sits between applied research and strong engineering. You will work on vision-language models, document processing, data curation, synthetic data generation, benchmarking, training, fine-tuning, and post-training. The goal is simple: make our document AI systems more accurate, faster, and more cost-effective in production.

You should be excited by frontier AI work, but equally motivated by practical product impact. This is not a pure research role where ideas stay in papers. You will be expected to prototype quickly, evaluate rigorously, and help turn promising approaches into production systems used by customers.

WHAT YOU’LL DO

- Develop and train vision-language models for document processing and document understanding.

- Build data pipelines for data curation, synthetic data generation, labeling, and benchmark creation.

- Evaluate base models and perform post-training or fine-tuning to hit specific performance targets.

- Improve model accuracy, latency, and cost-effectiveness across real-world document workflows.

- Design and maintain benchmarks to measure extraction quality, layout understanding, OCR performance, reasoning accuracy, and end-to-end system reliability.

- Work with messy real-world documents, including PDFs, scanned documents, tables, charts, forms, and multi-page enterprise documents.

- Collaborate with engineering to move successful research prototypes into production.

- Work directly with customers when needed to translate product requirements into benchmarks, experiments, and model improvements.

- Stay close to the latest research in vision-language models, document AI, post-training, synthetic data, and agentic systems.

- Use modern AI coding workflows and tools to move quickly.

WHAT WE’RE LOOKING FOR

- 3–7 years of experience in machine learning engineering, applied research, or research engineering.

- Strong ML foundation, including hands-on experience benchmarking and training models.

- Strong Python skills and comfort with modern ML tooling, especially PyTorch.

- Experience with computer vision, vision-language models, NLP, document AI, OCR, extraction, or agentic AI systems.

- Ability to build experiments, evaluate results, and iterate quickly toward measurable performance improvements.

- Strong engineering judgment and ability to write clean, production-quality code.

- Comfort working in a fast-paced startup environment with high ownership and limited structure.

- Adaptable, scrappy, and self-directed — someone who can figure things out without waiting to be told.

- Strong technical writing and communication skills.

NICE TO HAVE

- Prior startup experience, especially at an early-stage or high-growth AI company.

- Experience as a founder or early startup engineer.

- Experience building or improving document processing systems.

- Experience with synthetic data generation, post-training, fine-tuning, or benchmark design.

- Familiarity with tools such as vLLM, Pydantic, uv, ruff, mypy, Claude Code, Cursor, or similar modern AI engineering workflows.

- Experience with open-source AI infrastructure or developer tools.

WHO YOU’LL WORK WITH

You will work closely with the CTO and the document understanding team, partnering across research, engineering, product, and customer-facing teams.

WHY JOIN LLAMAINDEX

- Work on a core AI infrastructure problem: making complex documents understandable and actionable for AI systems.

- Build production systems at the frontier of vision-language models and document AI.

- Join a fast-growing startup with strong open-source adoption and commercial traction.

- Work directly with technical founders and a highly ambitious engineering team.

- Have real ownership over model quality, product capability, and technical direction.

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

LlamaIndex does not accept unsolicited agency resumes. Please do not forward resumes to our jobs alias, employees, or any other organization location. LlamaIndex is not responsible for any fees related to unsolicited resumes.

Work location

  • San Francisco, CA, US

Ready to make a decision?

This role is either worth your time or it isn’t.

Analyze the posting against your experience, see the gaps clearly, and build the right materials only if the opportunity makes sense.

Nothing is submitted automatically. You choose what happens next.

About this listing

Published on Ashby under the board identifier Llamaindex, which is the name the employer’s own job board carries. RoleSprint has not verified the company’s registered or trading name, so it is shown exactly as published rather than tidied up.

RoleSprint is not the employer and not a recruiter. Applications are made on the employer’s own site and never reach us; what RoleSprint does is help you decide whether a role is worth your time and prepare for it if it is.

Published 8 July 2026, last checked about a month ago. A posting stops being advertised here 90 days after the employer published it, and one the employer takes down is marked closed rather than quietly removed.

More searches like this one

Browse all current openings

No credit card requiredStart free