ABOUT MITHRL
We imagine a world where new medicines reach patients in months, not years, and where scientific breakthroughs happen at the speed of thought.
Mithrl is building the world’s first commercially available AI Co-Scientist. It is a discovery engine that transforms messy biological data into insights in minutes. Scientists ask questions in natural language, and Mithrl responds with analysis, novel targets, hypotheses, and patent-ready reports.
No coding. No waiting. No bioinformatics bottlenecks.
We are one of the fastest growing tech bio companies in the Bay Area with 12x year over year revenue growth. Our platform is used across three continents by leading biotechs and big pharmas. We power breakthroughs from early target discovery to mechanism-of-action. And we are just getting started.
ABOUT THE ROLE
We are hiring a Data Engineer, Knowledge Graphs to build the infrastructure that powers Mithrl’s biological knowledge layer. You will partner closely with the Data Scientist, Knowledge Graphs to take curated knowledge sources and transform them into scalable, reliable, production ready systems that serve the entire platform.
Your work includes building ETL pipelines for large biological datasets, designing schemas and storage models for graph structured data, and creating the API surfaces that allow ML engineers, application teams, and the AI Co-Scientist to query and use the knowledge graph efficiently. You will also own the reliability, performance, and versioning of knowledge graph infrastructure across releases.
This role is the bridge between biological knowledge ingestion and the high performance engineering systems that use it. If you enjoy working on data modeling, schema design, graph storage, ETL, and scalable infrastructure, this is an opportunity to have deep impact on the intelligence layer of Mithrl.
WHAT YOU WILL DO
- Build and maintain ETL pipelines for large public biological datasets and curated knowledge sources
- Design, implement, and evolve schemas and storage models for graph structured biological data
- Create efficient APIs and query surfaces that allow internal teams and AI systems to retrieve nodes, relationships, pathways, annotations, and graph analytics
- Partner closely with the Data Scientists to operationalize curated relationships, harmonized variable IDs, metadata standards, and ontology mappings
- Build data models that support multi tenant access, versioning, and reproducibility across releases
- Implement scalable storage and indexing strategies for high volume graph data
- Maintain data quality, validate data integrity, and build monitoring around ingestion and usage
- Work with ML engineers and application teams to ensure the knowledge graph infrastructure supports downstream reasoning, analysis, and discovery applications
- Support data warehousing, documentation, and API reliability
- Ensure performance, reliability, and uptime for knowledge graph services
WHAT YOU BRING
Required Qualifications
Strong experience as a data engineer or backend engineer working with data intensive systemsExperience building ETL or ELT pipelines for large structured or semi structured datasetsStrong understanding of database design, schema modeling, and data architectureExperience with graph data models or willingness to learn graph storage conceptsProficiency in Python or similar languages for data engineeringExperience designing and maintaining APIs for data accessUnderstanding of versioning, provenance, validation, and reproducibility in data systemsExperience with cloud infrastructure and modern data stack toolsStrong communication skills and ability to work closely with scientific and engineering teamsNice to Have
Experience with graph databases or graph query languagesExperience with biological or chemical data sourcesFamiliarity with ontologies, controlled vocabularies, and metadata standardsExperience with data warehousing and analytical storage formatsPrevious work in a tech bio company or scientific platform environmentWHAT YOU WILL LOVE AT MITHRL
You will build the core infrastructure that makes the biological knowledge graph fast, reliable, and usableTeam : Join a tight-knit, talent-dense team of engineers, scientists, and buildersCulture : We value consistency, clarity, and hard work. We solve hard problems through focused daily executionSpeed : We ship fast (2x / week) and improve continuously based on real user feedbackLocation : Beautiful SF office with a high-energy, in-person cultureBenefits : Comprehensive PPO health coverage through Anthem (medical, dental, and vision) + 401(k) with top-tier plans