Skip to main content
← Back to search
LS

Senior Data Engineer, Bioinformatics, Cheminformatics, Materials

Lila Sciences

$144,000 - $240,000 / year

San Francisco, CA USAFull-timeOn-siteProficient

Need a reasonable accommodation to apply or interview? Contact us.

JobMinglr uses automated technology to recommend jobs based on profile information and job preferences. Match Score does not determine eligibility for a position, prevent a user from viewing or applying to a job, or make hiring decisions on behalf of an employer.

Description

Your Impact at LILA

Lila’s mission is to accelerate scientific discovery with AI, and that depends on trustworthy scientific data. As a Data Engineer, you’ll build ETL pipelines and data models for Lila’s scientific data platform, working at the intersection of data engineering, computational biology, chemistry, and materials science.

You’ll partner with AI researchers and experimentalists to turn raw lab instrument outputs into validated, analysis-ready datasets. The core challenge is data modeling: transforming messy, per-instrument measurements into clean, well-typed data that is efficient to query, reliable to use, and ready for downstream analysis.

You’ll also build domain-specific analysis functions and reusable data pipelines that help scientists and AI researchers move faster without re-deriving bespoke solutions.

What You'll Be Building

• Design pipelines that turn raw lab output into analysis-ready scientific data. • Model heterogeneous data from bio, chemistry, and materials instruments. • Build validation checks, schema-evolution gates, and data quality workflows. • Develop reusable analysis functions for scientific and AI research workflows. • Improve automation and observability across instrument-to-result data flows. • Build canonical datasets that scientists and AI researchers can trust. • Use AI coding tools to accelerate pipeline development and team velocity.

What You'll Need to Succeed

• 2–6 years of experience in data engineering, bioinformatics, cheminformatics, or computational science. •
Strong Python skills, including typed, tested, production-quality code.
• Strong SQL skills, especially with Postgres or similar relational databases.
• Experience building ETL pipelines, data models, and reusable data transformations.
• Data science foundation, including statistics and pandas, NumPy, or similar tools.
• Experience translating noisy scientific measurements into accurate, validated datasets.
• Workflow orchestration experience, ideally Flyte, Airflow, Prefect, Dagster, or Nextflow.
• Active use of AI coding tools in day-to-day engineering work.

Bonus Points For

• Experience with columnar or lakehouse stacks such as Parquet, Iceberg, DuckDB, Polars, or Ibis.
• Familiarity with event-driven pipelines such as NATS or Kafka. • Exposure to lab instrument data formats, LIMS, or ELN systems.
• Familiarity with life sciences assays, sequencing, imaging, or flow cytometry.
• Familiarity with materials or chemistry methods such as XRD, XRF, SEM, TGA, or DSC.
• Experience with curve fitting, peak detection, or unit and dimensional analysis.

Compensation

We offer competitive base compensation with bonus potential and generous early-stage equity. Your final offer will reflect your background, expertise, and expected impact.

U.S. Benefits. Full-time U.S. employees receive a comprehensive benefits program including medical, dental, and vision coverage; employer-paid life and disability insurance; flexible time off with generous company wide holidays; paid parental leave; an educational assistance program; commuter benefits, including bike share memberships for office based employees; and a company subsidized lunch program.

International Benefits. Full-time employees outside the U.S. receive a comprehensive benefits program tailored to their region. USD salary ranges apply only to U.S.-based positions; international salaries are set to local market.

Expected Base Salary Range
$144,000—$240,000 USD

About LILA

Lila Sciences is building Scientific Superintelligence™ to solve humankind's greatest challenges. We believe science is the most inspiring frontier for AI. Rather than hard-coding expert knowledge into tools, LILA builds systems that can learn for themselves.

LILA combines advanced AI models with proprietary AI Science Factory™ instruments into an operating system for science that executes the entire scientific method autonomously, accelerating discovery at unprecedented speed, scale, and impact across medicine, materials, and energy. Learn more at www.lila.ai.

Guided by our core values of truth, trust, curiosity, grit, and velocity, we move with startup speed while tackling problems of historic importance. If this sounds like an environment you'd love to work in, even if you don't meet every qualification listed above, we encourage you to apply.

We’re All In

Lila Sciences is committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status.

Information you provide during your application process will be handled in accordance with our Candidate Privacy Policy.

A Note to Agencies

Lila Sciences does not accept unsolicited resumes from any source other than candidates. The submission of unsolicited resumes by recruitment or staffing agencies to Lila Sciences or its employees is strictly prohibited unless contacted directly by Lila Science’s internal Talent Acquisition team. Any resume submitted by an agency in the absence of a signed agreement will automatically become the property of Lila Sciences, and Lila Sciences will not owe any referral or other fees with respect thereto.

Benefits

  • health insurance
  • vision insurance
  • stock options
  • parental leave
  • disability insurance

About this role

Lila Sciences is building AI systems for scientific discovery, and this role sits at the core of that mission: you'll design and maintain the data pipelines that transform raw outputs from lab instruments—across biology, chemistry, and materials science—into clean, validated datasets that researchers and AI models can trust. The work is fundamentally about data modeling: taking messy, heterogeneous measurements and structuring them into efficient, well-typed data with strong validation and quality checks. You'll also build reusable analysis functions that let scientists and researchers move faster without reinventing solutions for each new experiment.

To succeed, you need solid data engineering fundamentals: 2–6 years of hands-on experience, strong Python and SQL (especially Postgres), and proven ability to build ETL pipelines and data models. You should have a data science foundation—comfort with statistics and libraries like pandas and NumPy—and real experience turning noisy scientific measurements into accurate, validated datasets. Workflow orchestration experience (Airflow, Flyte, Prefect, Dagster, or Nextflow) is expected, as is active use of AI coding tools in your day-to-day work.

Background in bioinformatics, cheminformatics, or computational science is valuable, and domain knowledge—familiarity with lab instruments, LIMS systems, sequencing, imaging, materials characterization methods, or specific assays—will set you apart. The role is full-time, in-office in San Francisco, with a base salary of $144,000–$240,000 plus bonus and early-stage equity, along with comprehensive U.S. benefits including medical, dental, vision, paid parental leave, and a subsidized lunch program.

How this employer is doing

solid

  • H1B: Sponsors Visas
  • Recently Raised Funding
  • Sponsors Visas

This role's local market on JobMinglr

Pay for this role

$144,000 to $240,000 per year

That is the employer's posted range, and where you land in it is usually decided in one conversation. Before that conversation, run the numbers: analyze an offer for this role.

How JobMinglr reads this job

Every listing here is scored against your profile before you apply: skills overlap, experience level, location and work arrangement, each weighted and explained. You see the score and the reasons, not just a list. How the matching works.