Agentic AI Data Engineer - CMC Data Integration

BioSpace - Indianapolis, IN

Hiring: Agentic AI Data Engineer - CMC Data Integration Company: BioSpace Location: Indianapolis, IN Job Posted Time: 2026-09-03 10:25:14 Employment Type: Full-time Target Skills & Keywords : AWS, Airflow, Azure, Azure DevOps, CI/CD, Data Mesh, Databricks, ELT, ETL, GitHub Actions, LLM, Lambda, LlamaIndex, MLOps, Prefect, Python, Redshift, S3, SQL About the job Experience: •2 years of relevant experience; OR BS in Computer Science or Computer Engineering with 3–5 years of hands-on data engineering experience. Required Skills: •Implement individual agent components (e.g., document extraction agent, schema mapping agent, validation agent) within the established orchestration framework (LangGraph, LlamaIndex, or equivalent) •Write tool-calling logic, handle failure modes, and ensure each agent component is testable and observable with instrumented logging of inputs, outputs, and intermediate decisions •Iterate on agent behavior based on real data performance; work with the senior engineer to identify and resolve failure patterns •Participate in validation and qualification activities for AI-assisted workflows, supporting documentation that demonstrates computational tools reflect scientific intent •Build review queues and flagging logic that surface low-confidence or out-of-specification extractions to scientific reviewers for approval before data is loaded •Implement routing logic that captures reviewer decisions, logs outcomes with full audit trail, and reintegrates approved data into the pipeline per 21 CFR Part 11 electronic records requirements •Tune flagging thresholds based on feedback from scientific owners; maintain and improve HITL logic as new data sources are onboarded •Design and build AI-assisted ingestion pipelines that extract and structure the data from unstructured CDMO/CRO data sources: PDFs (Certificates of Analysis, batch records), Excel files, and vendor portal exports Qualifications: •MS in Computer Science, Computer Engineering, Data Engineering, or related technical field with 1–2 years of relevant experience; OR BS in Computer Science or Computer Engineering with 3–5 years of hands-on data engineering experience. •Proficiency in Python and SQL; ability to write, review, and own production-quality code. •Demonstrated experience building ETL/ELT pipelines from unstructured or semi-structured sources (PDFs, Excel, JSON, XML). •Hands-on experience building LLM-powered applications: retrieval-augmented generation, tool-calling, multi-step orchestration, or equivalent agentic patterns. •Applied hands-on capability in cloud data platforms: Azure (Data Factory, Databricks, Fabric) or AWS (S3, Glue, Lambda, Redshift). •Solid understanding of relational data modeling, schema design, and data normalization principles. •Operational familiarity with data orchestration tools (Airflow, Azure Data Factory, Prefect, or similar). •Qualified applicants must be authorized to work in the United States on a full-time basis. Lilly will not provide support for or sponsor work authorization or visas for this role, including but not limited to F-1 CPT, F-1 OPT, F-1 STEM OPT, J-1, H-1B, TN, O-1, E-3, H-1B1, or L-1. •Solid functional working knowledge of 21 CFR Part 11, ALCOA+, and GxP data integrity principles, or clear demonstrated ability to apply similar audit/compliance frameworks. •Operational familiarity with pharmaceutical CMC data types: analytical results, batch records, stability studies, specifications. Compensation: •$65,250 - $169,400 / year •Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance) Interested candidates, please apply directly through the job posting on company's career page or try via AI auto apply on this platform. Don't miss this opportunity to join a forward-thinking team!