OpenAI

AI Research and Deployment

Researcher,Artifacts-AgentPost-Training

$250–380k San Francisco, California, United States FULL TIME Remote Friendly

Market Sentiment

HIGH DEMAND

Neural analysis suggests this role is
optimal for Mid+ candidates.

The Brief

“Researcher, Artifacts - Agent Post-Training at OpenAI. Skills: Agent Post-Training, Frontier Models, Artifact Creation. train frontier models. create polished work products”

What You'll Achieve.

create polished, useful work products; move from a vague user goal to a finished artifact; ship improvements into products used by real people; make agents genuinely useful

Industry & Context.

AI Research and Deployment

Problems you'll solve

move from vague behavioral problem to concrete experiment; define hypothesis; build pipeline; run model; analyze result; decide what to do next; debug hard failures

What They're Looking For.

Must Have

technical fundamentals in machine learning, technical fundamentals in software engineering, technical fundamentals in systems, technical fundamentals in statistics, hands-on experience with LLMs, hands-on experience with RL, hands-on experience with RLHF/RLAIF, hands-on experience with post-training, hands-on experience with evals, hands-on experience with graders, hands-on experience with synthetic data, hands-on experience with model training, hands-on experience with coding agents, hands-on experience with tool-using agents, hands-on experience with production ML systems

Nice to Have

prior background in consulting, prior background in finance, prior background in marketing, prior background in operations, prior background in data science

What You'll Do.

train frontier models

create polished work products

move from vague user goal to finished artifact

own improvements across post-training stack

build evals and environments

turn failures into training data

partner with product teams

translate product signal into model improvements

work on early-training interventions

shape downstream agent behavior

decide integrations for model runs

improve machinery for large-scale training

take on cross-functional projects

debug hard failures in models

run experiments that improve agentic model behavior

How You'll Work.

Team & Collaboration

work with researchers; work with engineers; work with product teams; work with infrastructure teams; work with safety/alignment partners; partner with Codex and ChatGPT product teams; comfortable working across research, product, infrastructure, data, evals, and safety boundaries; communicate clearly with each group

Communication Scope

communicate clearly with each group

Full Job Description

About the Team The Agent Post-Training team creates the frontier agents OpenAI ships to the world. We are training the models behind our agents in Codex, ChatGPT, the API, and other frontier products: persistent, proactive intelligence that can operate computers, collaborate with people and other agents, and expand what people and organizations can imagine, attempt, and achieve. We define what the next generation of agents should be able to do, build the training signal that teaches those abilities, and run the experiments that make them real. Our work spans coding, tool use, computer use, multi-agent coordination, long-horizon execution, factuality, instruction following, calibrated reasoning, and taste. Our team is where new model capabilities get made. We build the data, environments, graders, training methods, and feedback loops that shape what OpenAI's next agents can do, then carry those capabilities through major training runs and into the products people use. About the Role As a member of Agent Post-Training, Artifacts, you will train frontier models to create polished, useful work products: documents, spreadsheets, slide decks, dashboards, reports, analyses, and other interactive or editable artifacts. You will help teach our models to move from a vague user goal to a finished artifact with strong structure, visual taste, domain judgment, correctness, and low latency. This work will require owning improvements across our post-training stack, including RL, data pipelines, graders, reward signals, evals, and behavioral analysis. You will work with researchers, engineers, product teams, infrastructure teams, and safety/alignment partners to decide what should go into major model runs, measure whether it worked, and ship improvements into products used by real people. This is a high-agency role for people who want their work to land directly in frontier models. In this role, you will: - Design and run experiments that improve agentic model behavior for complex

Free ATS check

Applying for this Researcher, Artifacts - Agent Post-Training role?

Most applicants get filtered before a human reads their resume. See if yours makes the cut.

Should you apply? AI reads your resume vs this job — match score, gaps to address, ATS keywords.

SKILL SIGNAL 26 detected · ranked by frequency

Agent Post-Training ×2

Frontier Models ×2

Artifact Creation ×2

LLMs

RLHF/RLAIF

post-training

evals

graders

synthetic data

model training

coding agents

tool-using agents

production ML systems

Codex

ChatGPT

API

product impact

model behavior

research taste

engineering execution

agent usefulness

agent reliability

agent honesty

agent taste

agent ease of use

Role Details

Type FULL TIME

Category agents

Salary Band 200k+

AI-Extracted Insights

Domain Areas

persistent-intelligenceproactive-intelligenceoperate-computerscollaborate-with-peoplecollaborate-with-agentscodingtool-usecomputer-use

How to Apply on Ashby

Ashby is a fast modern ATS — most applications take under 3 minutes.
The resume parser is strong; verify parsed experience dates and job titles.
Custom screening questions are often scored algorithmically — answer completely.
Location field affects geo-based screening; use your actual metro area.

ANONYMOUS · UNFILTERED

What do employees actually say about OpenAI?

Real rants from real employees. Read before you apply.

Read Company Rants →