OpenAI

AI Research and Deployment

Researcher,ComputerUse-AgentPost-Training

$250–380k San Francisco, California, United States FULL TIME Remote Friendly

Market Sentiment

HIGH DEMAND

Neural analysis suggests this role is
optimal for Mid candidates.

The Brief

“Researcher, Computer Use - Agent Post-Training at OpenAI. Skills: Agent Post-Training, Computer Use, Model Training. Design experiments. Run experiments”

What You'll Achieve.

Ship frontier agents; Train models behind agents; Make agents persistent; Make agents proactive; Make agents operate computers; Make agents collaborate; Make agents expand imagination; Make agents expand attempts; Make agents expand achievements; Shape computer-use capabilities; Ship improvements into products; Make agents genuinely useful

Industry & Context.

AI Research and Deployment

Problems you'll solve

Reason through complex workflows; Debug hard failures; Turn messy qualitative behavior into concrete hypotheses

What They're Looking For.

Must Have

Technical fundamentals in machine learning, Technical fundamentals in software engineering, Technical fundamentals in systems, Technical fundamentals in statistics, Hands-on experience with LLMs, Hands-on experience with RL, Hands-on experience with RLHF/RLAIF, Hands-on experience with post-training, Hands-on experience with evals, Hands-on experience with graders, Hands-on experience with synthetic data, Hands-on experience with model training, Hands-on experience with coding agents, Hands-on experience with tool-using agents, Hands-on experience with production ML systems

Nice to Have

Experience with computer use, Experience with multi-agent coordination, Experience with long-horizon execution, Experience with factuality, Experience with instruction following, Experience with calibrated reasoning, Experience with browser navigation, Experience with desktop navigation, Experience with tool usage, Experience with complex workflows, Experience with user collaboration, Experience with long-horizon tasks, Experience with reliability, Experience with judgment

What You'll Do.

Improve agentic model behavior

Own improvements to post-training stack

Turn failures into training data

Translate product signal

Work on early-training interventions

Work on alignment interventions

Improve machinery for training

Improve machinery for launch

Take on cross-functional projects

Turn qualitative behavior into hypotheses

Turn qualitative behavior into experiments

Turn qualitative behavior into fixes

How You'll Work.

Team & Collaboration

Collaborate with people; Collaborate with other agents; Work with researchers; Work with engineers; Work with product teams; Work with infrastructure teams; Work with safety/alignment partners; Communicate clearly with each group

Communication Scope

Communicate clearly

Full Job Description

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Computer Use, you will teach models to operate computers. You will help train models that can navigate browsers and desktops, use tools and applications, reason through complex workflows, collaborate with users and other agents, and complete long-horizon tasks with reliability and judgment. This work sits at the intersection of frontier model training, product behavior, evaluation, and systems engineering, and will directly shape the computer-use capabilities shipped in OpenAI’s next generation of agents. Currently, our models are the best in the world at this behavior! You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you might - Design and run experiments

Free ATS check

Applying for this Researcher, Computer Use - Agent Post-Training role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

Should you apply? AI reads your resume vs this job — match score, gaps to address, ATS keywords.

SKILL SIGNAL 30 detected · ranked by frequency

Model Training ×6

Machine learning ×4

Software engineering ×4

Systems ×4

Statistics ×4

LLMs ×4

RL ×4

RLHF ×4

RLAIF ×4

Post-training ×4

Evals ×4

Graders ×4

Synthetic data ×4

Coding agents ×4

Tool-using agents ×4

Production ML systems ×4

Agent Post-Training ×2

Computer Use ×2

Codex ×2

ChatGPT ×2

API ×2

Research taste

Engineering execution

Product impact

Model behavior

Agent usefulness

Agent reliability

Agent honesty

Agent taste

Agent usability

Role Details

Experience 2–5 yrs

Level Mid

Type FULL TIME

Category agents

Salary Band 200k+

AI-Extracted Insights

Domain Areas

agent-trainingcomputer-usemulti-agent-coordinationlong-horizon-executionfactualityinstruction-followingcalibrated-reasoningtaste

How to Apply on Ashby

Ashby is a fast modern ATS — most applications take under 3 minutes.
The resume parser is strong; verify parsed experience dates and job titles.
Custom screening questions are often scored algorithmically — answer completely.
Location field affects geo-based screening; use your actual metro area.

ANONYMOUS · UNFILTERED

What do employees actually say about OpenAI?

Real rants from real employees. Read before you apply.

Read Company Rants →