Take the grind out of
the job hunt.
Let Talenry find matching jobs, tailor your applications, pitch you directly to hiring managers and fill your calendar with interviews.
Poke around. See Talenry in action.
1,316 open listings
Jobs Directory, the master index for your next role.
Showing AI engineer roles in the United States — E-Verified, paying $150k and up, posted within 7 days. Clear all
AI Engineer, Housing Data Systems
E-VerifiedNYU Furman Center — New York, NY
Posted 2 days ago
Saved
Assessment
- Overall match
- 92%
- Title fit
- 94%
- Skills similarity
- 91%
- Seniority
- Match
Strong match — Talenry can submit this one automatically if it can also fill the form confidently. Otherwise it pauses for your review.
Description
Own ingestion, chunking and embedding pipelines over heterogeneous municipal data, and the evaluation harness that decides whether a model change ships.
What you'll do
- Own ingestion, chunking and embedding pipelines over heterogeneous municipal data.
- Build the evaluation harness that decides whether a model change ships.
- Partner with researchers to turn questions into reproducible queries.
Requirements
- 3+ years building production Python data or ML systems.
- Hands-on with vector search and LLM evaluation.
- Comfortable owning correctness, not just throughput.
Product Engineer, Design Systems
E-VerifiedFigma — New York, NY
Posted 2 days ago
Saved
Assessment
- Overall match
- 88%
- Title fit
- 86%
- Skills similarity
- 90%
- Seniority
- Close
Good match on craft, a stretch on seniority. Talenry will draft it and wait for you.
Description
Build and maintain the component systems every Figma surface uses. TypeScript, React, an eye for typography and accessibility.
What you'll do
- Own the component systems every surface uses.
- Hold the typography and accessibility bar across teams.
- Ship the tooling designers live in daily.
Requirements
- Deep TypeScript and React.
- A portfolio showing systems work, not just screens.
- Comfort working directly with designers.
Research Engineer, Evaluation Infrastructure
E-VerifiedAnthropic — San Francisco, CA
Posted 3 days ago
Saved
Assessment
- Overall match
- 86%
- Title fit
- 90%
- Skills similarity
- 88%
- Seniority
- Stretch
A genuine stretch at lead level, but the evaluation experience is the exact thing they are hiring for.
Description
Own the evaluation infrastructure that decides when a model is ready: the harnesses, the datasets and the release gates research teams depend on.
What you'll do
- Own the harnesses, datasets and release gates research depends on.
- Decide when a model is ready to ship.
- Build the infrastructure other teams evaluate against.
Requirements
- Production evaluation infrastructure experience.
- Strong systems fundamentals.
- Willingness to own a release gate.
Research Engineer, Retrieval
E-VerifiedCohere — Remote (US)
Posted 5 days ago
Saved
Assessment
- Overall match
- 81%
- Title fit
- 82%
- Skills similarity
- 84%
- Seniority
- Match
Solid match and fully remote. Talenry will submit this automatically.
Description
Own retrieval quality end to end, from chunking strategy to the eval harness that catches regressions before customers do.
What you'll do
- Own retrieval quality end to end.
- Set chunking strategy and measure it.
- Catch regressions before customers do.
Requirements
- Retrieval and embedding experience.
- Comfort with evaluation harnesses.
- Remote-first working.
Applied AI Engineer
Not enrolledRamp — New York, NY
Posted 8 days ago
Saved
Assessment
- Overall match
- 64%
- Title fit
- 66%
- Skills similarity
- 68%
- Seniority
- Match
Weaker match, and the employer is not enrolled in E-Verify. Worth a look only if you are flexible on sponsorship.
Description
Build the models that read receipts, catch duplicate spend and draft accounting entries.
What you'll do
- Build the models that read receipts and catch duplicate spend.
- Draft accounting entries automatically.
- Work close to the finance product surface.
Requirements
- Applied ML in production.
- Comfort with messy financial documents.
- Product instincts.
Sarah Johnson
New York, NY | linkedin.com/in/sarahjohnson | github.com/SarahJohnson
Professional summary
AI Engineer with deep experience evaluating and shipping large language model systems. Builds the data pipelines and human-in-the-loop review that turn research prototypes into production tools.
Experience
GoogleManhattan
AI ResearcherSep 2025 – Present
- Engineered complex reasoning evaluation tasks for LLMs in Python and PyTorch, increasing measured accuracy 20% and surfacing failure modes via adversarial agent simulation.
- Architected agent responses to ML tasks, improving rogue-behaviour detection by 30%.
- Led a team of 10 researchers on model-limitation studies that reshaped internal AI tooling.
AppleManhattan
AI Systems EngineerJun 2023 – Sep 2025
- Designed and deployed NLP and speech models for intent recognition, cutting query processing time 25%.
- Optimised on-device inference with TensorFlow Lite — 40% lower latency at 95% accuracy.
Projects
- Housing-data RAG pipeline — ingest, chunk and embed 1.2M municipal records; 90% retrieval precision at p95 under 400 ms.
Skills
Python, SQL, Go, TypeScript, PyTorch, TensorFlow, Hugging Face Transformers, LangGraph, FastAPI, Postgres/pgvector, MLflow, CI/CD
The posting asks for someone who owns ingestion over heterogeneous municipal data and the harness that decides whether a change ships. That is the job I have been doing at Google, on a different corpus.
I built the evaluation harness for multi-step reasoning failures — deliberately not aggregate accuracy, because aggregate accuracy hides the mechanism.
On heterogeneous sources specifically: I would normalise every source onto a canonical entity before embedding, so the eval set stays stable while the sources churn.
Tailor my resume for the NYU Furman Center AI Engineer role
Pulled the posting. Strong fit — here's the match before I draft.
AI Engineer, Housing Data Systems↗
NYU Furman Center · New York, NY
Match92%
Salary · Type$165k – $195k · Remote
Rewrote the summary around LLM evaluation, promoted the Google reasoning-eval bullet, and swapped in the RAG pipeline project — it maps to their data-ingest requirement.
Resume & cover letter↗
AI Engineer · NYU Furman Center
Both documents ready● Open
● Live⤓ Download
Make the summary two lines shorter
Done — page 1 re-rendered. The preview on the left is already updated.
Auto apply
Your applications, on autopilot.
Auto-apply to the roles that clear both bars.
A job applies on its own only when it clears both bars — a strong match and a form we can complete. Everything else waits here for your review.
- Filed this month26
- Avg match84%
- Time saved11h
- Monthly quota26/ 500
The agent stopped at something only you can answer. Nothing sends itself from here.
- AI Engineer, Housing Data Systems — NYU Furman Center
- Senior Applied Scientist, Search Ranking — Amazon Answer
- Research Engineer, Retrieval — Cohere Watch
| Role | ATS | When | Match | State |
|---|---|---|---|---|
| Staff Machine Learning EngineerAnthropic | greenhouse | 1h ago | 88% | Submitted |
| LLM Platform EngineerDatabricks | lever | 3h ago | 81% | Submitted |
| Product Engineer, Design SystemsFigma | greenhouse | Yesterday | 83% | Awaiting confirmation |
| ML Platform EngineerSpotify | lever | Yesterday | 80% | Submitted |
Reach out to the hiring manager for the NYU Furman Center role directly.
Drafted a note that leads with the evaluation-harness work — it is the closest match to what they listed first.
To Hiring Manager · NYU Furman CenterSent · 2m ago
I saw the AI Engineer opening on the housing-data team. I built the evaluation harness for multi-step reasoning failures at Google — the same shape of problem as your ingestion pipeline — and would welcome the chance to talk it through.
They replied — asking about your availability for a call this week.
From Hiring Manager · NYU Furman Center
Thanks for reaching out directly — this is exactly the background we are looking for. Do you have 20 minutes Thursday or Friday?
From Recruiter · Cohere
Saw your profile through Talenry — we have a Retrieval role that looks like a strong match. Do you have time this week for a quick call?
A recruiter from Cohere reached out about the Retrieval role — 81% match. Want me to reply?
Yes — say I am interested and free Thursday afternoon.
Sent. I mentioned your retrieval-quality work from the evaluation harness to keep the thread warm.
To Recruiter · CohereSent · just now
Thanks for reaching out — I am interested and free Thursday afternoon. Happy to talk through the retrieval-quality work I have been doing, which lines up closely with this role.
← Back to dashboard
Run a mock interview for the Staff AI Engineer role at Anthropic.
Done — scored 74. Full report below.
Interview Report
Staff AI Engineer @ Anthropic · Mar 12, 2026, 4:36 PM · 8:12
74/100 Overall score- Technical depth 80/100
- Communication 78/100
- Confidence 64/100
- Role fit 81/100
- Authenticity 67/100
You are strongest when you are specific. The evaluation-harness answer landed because you named what you measured and why aggregate accuracy was the wrong metric. You lost ground on the ambiguity question — you jumped to a solution before checking what the interviewer actually meant by "meaningful". Slow down and restate the problem first.
Key strengths
- Concrete, numeric evidence for every claim. Never "improved performance".
- Clear separation of retrieval and reasoning failure modes; shows real evaluation instinct.
- Directly connected your latency work to their deployment constraints.
Areas for improvement
- Ask a clarifying question before designing. Twice you solved a problem that was not asked.
- Trailing "…does that make sense?" undercuts otherwise strong answers.
- No questions for the interviewer at the end. Prepare two about their data-quality process.
Question-by-question
Walk me through the evaluation harness you built at Google.88
You said: Measured multi-step reasoning failures via adversarial task variants, not aggregate accuracy.
Excellent. You led with the decision — what to measure — rather than the tooling, and the accuracy figure was tied to a mechanism. This is the answer to reuse.
How would you approach ingestion so evaluation stays meaningful?64
You said: Normalise every source onto a canonical entity before embedding.
The design is sound but you never asked what "meaningful" meant to them. One clarifying question would have doubled the value of this answer.
Morning, Sarah.
Ask for anything across your search. I have context on every role and document.
Try
Good morning, Sarah.
Three things are waiting on you, the most urgent being the NYU Furman Center application that's been ready to send since 4:12 PM yesterday. Overnight you picked up an interview invite from Google DeepMind and a reply from Cohere; Stripe passed. Seven new roles cleared your bar while you slept.
Nothing here moves without you.
The NYU Furman draft has all nine fields filled and is sitting on the approval gate at 92% match.
The Amazon run stalled on a single question it could not answer from your profile: "years of experience with learning-to-rank?" Answer
Thursday's DeepMind interview is in two days and you have not prepared for it. Your last mock flagged that you design before clarifying. Worth one run. 10 min
Overnight & yesterday
DeepMind moved to In process — your sixth interview — and Cohere's recruiter replied. Stripe closed you out on the Senior ML Engineer role. Behind the scenes, the agent applied to 6 roles and read 14 messages to keep the board honest. You ran one mock interview and scored 74.
Month to date: 6 interviews from 31 tracked applications (19%) — on 34 auto-applies, 84% average match.
Posted since yesterday, ranked by fit.
- 1 Machine Learning Engineer, Civic Data · Bloomberg Closest thing to the Furman role you've seen; same municipal-data problem. 91% · $170k
- 2 Applied Scientist, Ranking · Spotify Remote, E-Verified. Weaker on evaluation, strong on scale. 85% · $185k
- 3 Research Engineer, Evaluation · Scale AI Your exact title. No E-Verify record, which matters for your status. 83% · $200k
4 more between 61% and 78%See all
Thursday, 12 March
Hey Sarah, you're moving.
Your search stands at 6 interviews and one offer from 31 applications to date. 12 roles are active in your pipeline right now, and your nearest date is Thursday's DeepMind interview.
- Interviews landed6
- Interview rate19%
- Offers1
Across 31 applications to date
- Saved 5
- Applied 6
- In process 1
| Role | Match | Salary | Status | |
|---|---|---|---|---|
| AI Engineer, Housing Data SystemsNYU Furman Center · E-Verified | 92% | $165k–$195k · Remote | Docs ready | |
| Staff Machine Learning EngineerAnthropic · E-Verified | 88% | $240k–$340k · San Francisco | No docs yet | |
| Research Engineer, RetrievalCohere | 79% | $190k–$240k · Remote | No docs yet | |
| Senior ML EngineerStripe | 76% | $210k–$280k · Remote | No docs yet |
Roles
- Research Engineer, Evaluation InfrastructureAnthropic · 86 match
- AI Engineer, Housing Data SystemsNYU Furman Center · 92 match
Your documents
- Resume — Anthropic, Evaluation InfrastructureTailored 4 days ago
Conversations
- Staff MLE · AnthropicYou asked about the practice screen
AI Engineer · NYU Furman Center
Is this one worth applying to?
Yes — it is your strongest match at 92. The evaluation-harness responsibility is the thing you already did at Google, and they are E-Verified so sponsorship is not a dead end.
Go ahead and apply.
Submitted. I led the resume with the harness work and cut the latency bullet to make room. Nothing in the form committed you to a start date, so I did not need to stop.
Staff MLE · Anthropic
How did the practice screen go?
Seventy-four. Your harness answer was the strongest thing in it. You lost ground by designing before clarifying — twice you answered a question that had not been asked.
What should I do differently?
Ask one clarifying question before you design. And prepare two questions for them about their data-quality process; you had none at the end.
Six things doing work
while you are not looking.
-
Work authorisation, answered before you apply
Every listing carries its employer’s E-Verify status, and you can filter on it alongside visa sponsorship. If your right to work depends on the answer, finding out after four interviews is not a small inconvenience.
E-VerifiedNot enrolled
-
It finishes the forms the big boards put in the way
Most roles are gated by an applicant tracking system rather than an inbox. Talenry completes and submits on the ones that matter, including the multi-page screeners, and pauses when a question is genuinely yours to answer.
WorkdayGreenhouseLeverAshbyiCIMS
-
Your profile, in front of every company hiring here
Talenry works the company side too. Come in through us and your structured profile is visible to every employer running a search on the platform — so roles reach you that you never applied to, and you arrive already scored.
Nothing is shared until you say so, and you decide who can see you.
-
Rehearse the screen you are about to sit
Mock interviews are generated against the role you actually applied to, so you practise the questions that screen is likely to ask. Afterwards you get a scorecard, per question, and the specific thing to do differently.
-
An agent acting for you is one you can audit
Every application, message and answer sent in your name is visible to you. Anything that commits you to something waits for your approval, and nothing about your history is invented or embellished.
-
The board keeps itself current, in your inbox
Replies, rejections and interview invitations are read as they arrive and moved to the right column on their own. There is no spreadsheet to maintain and nothing to mark off by hand.
Free to search.
Paid when you want the agent working.
Annual billing saves 20%. Cancel any time.
The things people ask
before they trust it.
Does Talenry apply to things without asking me?
It applies within the limits you set, and anything that commits you to something waits for your approval. Every message sent in your name is visible to you, and you can stop the agent at any point.
Is anything invented about me?
No. Talenry rewrites and reframes what is already in your profile against the posting — it does not add experience, employers or credentials you did not give it. The record of what it did is yours to audit.
What does the E-Verify filter actually do?
It shows whether an employer is enrolled in E-Verify, and lets you filter on it alongside visa sponsorship. That tells you before you apply whether a role is realistically open to you.
Do I need a paid plan to use it?
No. Free covers search, match scores, E-Verify and visa filters, the inbox tracker, five tailored documents a month and three auto-applies. Paid plans exist for when you want the agent working at volume.
Can I cancel, and is annual cheaper?
Plans are monthly and you can cancel any time. Annual billing saves 20% — $15, $31 and $63 a month for Resume, Auto-Pilot and Power respectively.