All open positions

AI Safety Researcher I

MindBench.ai AI Evaluation Platform

Division of Digital Psychiatry · Beth Israel Deaconess Medical Center · Harvard Medical School · in partnership with NAMI

Full-Time Boston, MA (100% in person) Early-career

Position Overview

MindBench.ai is a joint effort between the Division of Digital Psychiatry at BIDMC, a Harvard Medical School teaching hospital, and the National Alliance on Mental Illness. We independently build and run evaluations of AI systems in mental health: safety benchmarks, multi-turn evals, audits of treatment and information preferences, and a public alignment standard. Clinicians, people with lived experience, and researchers contribute to every one of these efforts. Our results are published openly and transparently.

This is an early-career role in an academic lab based in Boston, fully in person.

What We Offer

  • Project ownership, early. Within a few months you will run a piece of the evaluation program end to end, from design through data collection, review, and publication
  • Weekly time with the lead engineer, plus training in the coding-agent workflows (Claude Code, Codex) we use across the platform
  • Authorship. Our evaluation work goes into peer-reviewed papers, and contributors are on them
  • Direct contact with a national network of clinicians, advocates, and industry partners who collaborate with us across projects
  • A Harvard Medical School-affiliated hospital and NAMI on your resume, alongside professional development opportunities
  • A hand in shaping how AI in mental health gets measured, alongside a collaborative team of researchers

What You Will Work On

Our priorities move with the field, so this is a sample rather than a definitive list. Nobody is expected to do all of this. We divide work based on people's strengths and interests, so you would have the flexibility to focus on the areas you are good at and most excited to work on.

If you build things:

  • Python pipelines that run models against our benchmarks and capture everything about each run
  • Platform work in TypeScript (Express, Prisma, React) on the system that stores, validates, and publishes results
  • Getting our coding-agent workflows to do more of the routine work safely

If you write and review:

  • Drafting and editing the scenarios, prompts, and rubrics that make up our evaluations
  • Reviewing model-generated content for realism, clinical plausibility, and quality before it reaches clinicians
  • Writing up findings for both research and public audiences

If you coordinate:

  • Working with volunteer clinicians and lived-experience reviewers, from recruitment through data collection
  • Keeping our external partners, steering committee, and collaborating labs informed and on schedule

Required Qualifications

  • Enthusiasm and interest in evaluating AI in mental health specifically, not only AI or mental health broadly: the particular question of whether these systems are safe or useful for people in distress, and how you would find out
  • A good and clear writer. We have lots of writing tasks!
  • Comfort with priorities that change. We are small and the field moves fast, so you must be flexible and enjoy a fast pace

Nice to Have

None of these are required:

  • All degrees and backgrounds are welcome to apply: psychology, medicine, CS, or no degree in particular. English literature PhDs, philosophy PhDs, former clinicians, self-taught engineers, psychometricians, and recent graduates are all welcome candidates for this role
  • Prior AI safety or machine learning experience
  • Python or coding-agent experience. A plus, since there is always more platform work, but a strong writer or coordinator is just as valuable to us right now

Details

  • Full-time position based at BIDMC in Boston, MA
  • BIDMC benefits, including health insurance, retirement, and hospital-employee benefits

About the Division

The Division of Digital Psychiatry advances mental health care through smartphone-based digital phenotyping, digital therapeutics, AI safety evaluation, digital literacy initiatives, and technology-enhanced clinical programs, including the Digital Clinic model for depression and anxiety. The Division's work has informed Congress, the FDA, the UK MHRA, the WHO, and the American Psychiatric Association.

Learn more: docs.lamp.digital

To Apply

We would love to learn more about you. Send the following to jtorous@bidmc.harvard.edu and mflather@bidmc.harvard.edu:

  1. A resume
  2. One paragraph on this: if you could build your own AI and mental health evaluation, what would you want to measure and why?
Apply by Email

Don't see a fit? We also welcome students, clinicians, researchers, and volunteers through our Join Us page.