Applied Research Scientist

Remote Full-time

About the Role We are seeking an Applied Research Scientist to design and run rigorous experiments for LLM-based agents, with a focus on clinical and agentic reliability. This role will own the development of automated evaluation frameworks, bridging research prototypes into production systems, and partnering closely with product and engineering to ensure safety, robustness, and measurable impact. About Us Team from OpenAI, DeepMind, NASA, GoogleX, Tesla, and 2 physicians: 6 exits, 2 IPOs. Our model outperforms Claude, Gemini, and GPT-4.5 on clinical benchmarks. 400+ healthcare orgs signed in 16 months. ⚡ $25M raised from YC, Amity Ventures, Sequoia scouts, and more. $1T+ market opportunity. We’re going after all of it. Key Responsibilities • Design and run experiments to measure accuracy, robustness, and hallucination rates in LLM agents. • Build automated evaluation pipelines (LLM-as-judge + human review) with clinical-grade benchmarks. • Partner with Research Ops/IRB to design efficacy studies and align with regulatory requirements. • Translate research into production-ready evaluation systems, collaborating with engineering to land features 0→1. • Develop error taxonomies, ablations, and guardrails to ensure safe and reliable agent behaviors. Hard Requirements • Proven experience designing agentic processes and LLM evaluation/benchmarking frameworks. • Strong Python and ML background (PyTorch/TensorFlow, Hugging Face, LangChain/LlamaIndex). • Demonstrated ability to design rigorous experiments and translate findings into production. • Track record of published research or deep applied work in LLMs and agent evaluation. • Strong communication and technical writing skills to articulate complex findings clearly. Nice-to-Have • Prior work in healthcare/clinical NLP with awareness of medical data standards. • Experience running IRB-aligned or clinical-grade studies. • Exposure to noisy/limited medical data and designing strategies to overcome constraints. First-Month Focus • Audit existing evaluation approaches for clinical and agentic tasks. • Define initial benchmarks and build early automated pipelines. • Partner with engineering to land first set of CI gates for accuracy, factuality, and safety. Success OKRs (90 Days) • Deliver a repeatable evaluation framework with automated pipelines in production. • Demonstrate measurable improvements in robustness, hallucination reduction, or safety. • Publish or present internal research findings that directly shape product reliability. Culture Fit • Persistent, driven problem solver • Willing to push back on leadership to defend quality/timelines • Thrives in high-ambiguity, fast-paced startup environments Why Join Sully.ai? Shape the Future of Healthcare: Build category-defining partnerships that enable doctors to focus on saving lives. Early-Stage Impact: Join early and play a critical role in shaping our partnership roadmap and overall company growth. Remote-First Culture: Work with a talented, mission-driven team in a flexible, remote environment. Competitive Compensation: Enjoy a competitive salary, equity, and the opportunity to make a real difference. Solve Scalability Challenges: Tackle complex challenges in a rapidly growing company, driving impactful change in healthcare. Sully.ai is an equal opportunity employer. In addition to EEO being the law, it is a policy that is fully consistent with our principles. All qualified applicants will receive consideration for employment without regard to status as a protected veteran or a qualified individual with a disability, or other protected status such as race, religion, color, national origin, sex, sexual orientation, gender identity, genetic information, pregnancy or age. Sully.ai prohibits any form of workplace harassment. Apply tot his job

Apply Now →

Experienced Customer Service Representative – Bilingual English/Spanish Preferred, Remote Work from Home Opportunity After Comprehensive Training

Remote Full-time

Applied Research Scientist

Similar Jobs

Flexible Part-Time Research Contributor (Hiring Immediately)

Academic Research Assistant - AI Trainer

Focus Group - online Research - High pay with flexible hours (Hiring Immediately)

Clinical Research Associate, Obesity/Diabetes/GLP-1 (Full Service) - IQVIA

Remote Bookkeeping & Accounting Specialist - Part-Time Opportunity

Clinical Informaticist - Optime/Anesthesia - IT-Clinical

Remote Property Management Assistant

Expert Contributor - Commercial Property & Facility Management

[Remote] Projects Operations Coordinator

FLEX Director, Global Property Management Systems - Daylight PMS

Experienced Customer Support Representative – Insurance Industry Expertise (Fully Remote)

Experienced Customer Service Representative – Work From Home Opportunity with arenaflex

Experienced Manager, Customer Success – Driving Growth and Retention for Ecommerce Brands at Blithequark

Experienced Customer Service Representative – Bilingual English/Spanish Preferred, Remote Work from Home Opportunity After Comprehensive Training

Prior Authorization (full Remote)

Patent Associate – Mechanical or Chemical Engineering (Minneapolis)

Experienced Remote Customer Support Specialist for Innovative Electric Vehicle Manufacturer – arenaflex

Experienced Customer Service Representative – Temporary Work-From-Home Opportunity with arenaflex

Amazon Customer Service Representative - Work From Home Opportunity with Competitive Pay ($16-$35/hr)

Staff Software Engineer, DevOps

Applied Research Scientist

Similar Jobs

Flexible Part-Time Research Contributor (Hiring Immediately)

Academic Research Assistant - AI Trainer

Focus Group - online Research - High pay with flexible hours (Hiring Immediately)

Clinical Research Associate, Obesity/Diabetes/GLP-1 (Full Service) - IQVIA

Remote Bookkeeping & Accounting Specialist - Part-Time Opportunity

Clinical Informaticist - Optime/Anesthesia - IT-Clinical

Remote Property Management Assistant

Expert Contributor - Commercial Property & Facility Management

[Remote] Projects Operations Coordinator

FLEX Director, Global Property Management Systems - Daylight PMS

**Experienced Customer Support Representative – Insurance Industry Expertise (Fully Remote)**

**Experienced Customer Service Representative – Work From Home Opportunity with arenaflex**

**Experienced Manager, Customer Success – Driving Growth and Retention for Ecommerce Brands at Blithequark**

Experienced Customer Service Representative – Bilingual English/Spanish Preferred, Remote Work from Home Opportunity After Comprehensive Training

Prior Authorization (full Remote)

Patent Associate – Mechanical or Chemical Engineering (Minneapolis)

Experienced Remote Customer Support Specialist for Innovative Electric Vehicle Manufacturer – arenaflex

Experienced Customer Service Representative – Temporary Work-From-Home Opportunity with arenaflex

Amazon Customer Service Representative - Work From Home Opportunity with Competitive Pay ($16-$35/hr)

Staff Software Engineer, DevOps

Experienced Customer Support Representative – Insurance Industry Expertise (Fully Remote)

Experienced Customer Service Representative – Work From Home Opportunity with arenaflex

Experienced Manager, Customer Success – Driving Growth and Retention for Ecommerce Brands at Blithequark