About the Role
Our client — a physician-founded, venture-backed clinical AI research company (client identity confidential) — is building a system that compiles the standard of care into decision artifacts that execute at the point of care: schema-enforced, verified for exhaustiveness, and traceable to the source evidence that justifies them. Their compiler is live in public health care today, with clinicians acting on its recommendations daily.
This is a research role with production stakes. You will study how far frontier models can be pushed in extracting, structuring, and proving clinical logic, design the experiments that settle what works, and ship the winning approach into a live pipeline. You'll feel at home if you have published in clinical AI, want out of the demo cycle, and believe the logic behind a clinical decision has to be inspectable.
What You'll Do
Build the pipeline — convert clinical standards into executable modules: parsing with document and layout models, transformation, and validation.
Design the target syntax for representing guideline logic (exclusions, precedence, recommendations).
Run the experiments — open vs. closed models, fine-tuning, constrained decoding.
Co-design benchmarks — datasets, metrics, and error analysis for clinical fidelity.
Build adversarial cases that measure source fidelity, logical completeness, and recommendation correctness.
Develop deterministic validators using static analysis, constraint solving, or formal methods to make ambiguity, contradiction, and source silence explicit.
Turn model failures into hypotheses, experiments, and better system design.
Work directly with the founders and the clinicians who stake their license on the outputs.
Requirements (Must-Have)
Experience working in a frontier AI lab.
3+ years of research or industry experience, or a final-year PhD/postdoc.
Shipped ML/NLP systems to production — not just research prototypes.
BS or higher in CS, ML, math, or a related technical field.
Production-level Python, PyTorch, and the Transformers ecosystem.
Experience with LLM evaluation, training, or prompt engineering.
Able to articulate failed experiments and uncertainty clearly in writing.
Willing to attend in-person offsites multiple times per year.
Authorized to work in the US or Canada — no visa sponsorship (potential exceptions for current employees of frontier labs).
Nice-to-Have
Compilers, program synthesis, formal methods, or static analysis.
Healthcare, safety-critical, or highly regulated domain experience.
PhD in CS, ML, NLP, formal methods, or statistics.
Knowledge of IR design, constraint solving, SAT/SMT, or DSLs.
Details
Compensation: USD $150K–$300K (higher possible for exceptional candidates), plus competitive equity (cash/equity mix is flexible).
Location: Fully remote, US or Canada, with offsites several times per year.
Type: Full-time.
Tech stack: Python, PyTorch, Transformers, LLMs, static analysis, SMT/SAT solvers, FHIR, HL7 v2.
A physician-founded, venture-backed clinical AI research company that builds systems to compile medical standards of care into executable decision artifacts deployed at the point of care. Their proprietary compiler converts clinical guidelines into schema-enforced, verified modules that execute in live healthcare settings, with clinicians actively using recommendations daily. The company combines frontier AI research with production stakes, focusing on extracting, structuring, and proving clinical logic through advanced language models and formal verification methods. Their mission is to make clinical decision-making more transparent, evidence-based, and inspectable through AI-powered systems that transform unstructured medical standards into actionable, verifiable recommendations.