All open roles

Member of Technical Staff, Agents

๐Ÿ“ Palo Alto, CA ยท Full-time ยท In office Monday, Wednesday, and Friday

Apply on LinkedIn

About Noah

Most AI assistants are copilots. They wait for you to type. Noah is an autonomous assistant for executives that does real work in the background, over text and voice. It books meetings, emails clients, makes phone calls, and keeps important relationships warm.

Noah launched as #1 Product of the Day, Week, and Month on Product Hunt. We're a small, senior team in Palo Alto, backed by Hustle Fund and Neon. Our founder, Ashish Toshniwal, bootstrapped his last company, YML, to $100M in annual revenue and a $350M exit. Engineering is led by our CTO, Ryan Brandt. We think we win on distribution, focus, and speed.

Read Ashish's launch post: https://www.linkedin.com/posts/ashishtoshniwal_i-have-been-preparing-for-this-moment-for-activity-7440049408421027840--EMM

The Role

We're hiring two engineers.

When Noah books a meeting or emails a client, it has to be right nearly every time. There is very little room for error, and the hardest part is the last 1%: the long tail of situations nobody has seen yet. Getting there takes serious evaluation work, careful harness engineering, and sometimes fine-tuning and post-training.

You'll build what Noah can do, and prove that it works. You'll design new capabilities, find where Noah fails in production, work out why, and fix the system rather than patching individual examples. Every change is proven on real conversations before it reaches users. Some people lean toward evals and some toward building. Here, you do both.

This is not primarily a prompt-engineering role or a narrowly scoped backend role. We work hard and we move fast. You'll own a whole area from day one, work directly with our CTO, and see your work reach executives the same week.

Tech stack: Python | Django | React | TypeScript | Custom Python agent harness

What You'll Work On

  • Agents that finish the job. Done means an executive hands off real work and it gets done end to end, without them checking.
  • Make "right nearly every time" measurable. Done means that for anything the agent does, we know how often it gets it right, and we trust that number.
  • New kinds of work, proven before launch. Done means the agent takes on tasks it couldn't do before, in the tools executives already use, and the evidence says it's ready before anyone relies on it.
  • Turn failures into fixes that stay fixed. Done means a failure in production becomes a proven fix quickly and never comes back.
  • Find the unknown unknowns. Done means you own detection at scale, using tools like Braintrust and Raindrop to catch both the failure patterns we already know and the long tail we haven't seen yet, before an executive notices.
  • Recovering on its own. Done means that when a tool or site fails, the agent still gets the job done.
  • Make agent improvement a repeatable science. Done means any engineer can change the agent, measure the effect, and know whether it's better.

You Might Be a Good Fit If

  • You come from a deeply agentic background. You've built or run agents that do real work for real users, and you care about shipping high-performing systems at scale that delight them.
  • You're an engineer first. You write maintainable production Python and strong SQL, and you're comfortable in Django, React, and TypeScript when the work crosses into them.
  • You love evals. You know the work of Hamel Husain and others on high-quality AI evals, and you know enough statistics and experiment design to tell a real improvement from noise.
  • You treat harness engineering as a repeatable science: change one thing, measure it with realistic end-to-end simulations, and keep what works. You know how different models behave and when to use each.
  • You can move across prompts, traces, model outputs, code, databases, and APIs to find the real cause of a failure, and you check that a number is true before you trust it.
  • You've owned something end to end, shipped it to real users, and can show what moved because of it.
  • You've stayed somewhere long enough to own your decisions and live with the results.

Strong Plus

  • Post-training, fine-tuning, or RL experience.

Experience with a particular agent framework is not required. Our agent harness is custom, and we care more about your ability to understand and improve the underlying system.

We encourage you to apply even if you don't meet every point above.

How We Hire

We interview on a real engineering problem we've already solved, and we decide fast.

Equal Opportunity

HeyNoah considers applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other characteristic protected under applicable law. We are committed to providing reasonable accommodations throughout the interview and hiring process.