Career CategoryInformation Systems
Job DescriptionPrincipalArchitect
What you will do
Let9s do this. Let9s change the world. In this vital role you will play a pivotal role in building and scaling our machine learning models from development to production. Yourexpertisein both machine learning and operations will be essential in creating efficient and reliable ML pipelines.A background in data engineering, including experience with data pipelines and distributed data processing, is a strong plus.
- Lead the end-to-end design, development, and delivery of machine learning and Generative AI (GenAI) solutions, leveraging Databricks, Apache Spark, SQL, and Python for scalable data processing, feature engineering, and model development from problem framing to production deployment and business impact realization.
- Act as an Architect for large-scale Data Engineering and ML/GenAI initiatives, driving architecture decisions across lakehouse platforms (Databricks), distributed compute (Spark), and cloud ecosystems (AWS/GCP/Azure) to ensure scalability, reliability, and long-term maintainability.
- Design and implement advanced data pipelines and AI systems, including batch and streaming data processing (Spark), data modeling (SQL), and ML workflows (Python), along with multi-agent architectures, reasoning workflows, tool integration, and autonomous decision-making systems.
- Build and optimize robust data foundations for AI by developing high-quality, scalable ETL/ELT pipelines in Databricks, ensuring data availability, consistency, and performance for downstream ML/GenAI use cases.
- Define and institutionalize evaluation, validation, and governance frameworks for ML/GenAI systems, including model performance tracking, prompt evaluation, safety guardrails, hallucination mitigation, and compliance.
- Partner directly with business stakeholders and product leaders to translate objectives into data-driven AI/ML solutions, ensuring measurable value through well-defined data pipelines, KPIs, and experimentation frameworks.
- Establish and enforce best practices in MLOps, LLMOps, DataOps, and DevOps, including CI/CD pipelines, Databricks workflows, monitoring, observability, reproducibility, and cost optimization.
- Architect and oversee scalable cloud-based data and AI platforms, integrating Databricks Lakehouse, Spark processing layers, and cloud-native services for unified analytics and AI workloads.
- Drive experimentation strategy, including A/B testing, prompt optimization, and data-driven iteration, leveraging SQL analytics and Python-based experimentation frameworks.
- Provide mentorship to L4 and L5 engineers in data engineering (Spark, SQL, Databricks) and AI/ML development (Python, GenAI frameworks), including design reviews, code reviews, and career guidance.
- Lead cross-functional collaboration across data engineering, data science, platform engineering, and business teams to deliver integrated, production-grade AI solutions.
- Stay at the forefront of advancements in data engineering (Spark ecosystem, lakehouse architectures) and Generative AI/agentic systems, driving adoption of new technologies and best practices.
What we expect of you
We are all different, yet we all use our unique contributions to serve patients. The professional we seek is a Software Engineer with these qualifications.
Basic Qualifications
- Doctorate degree and 2 years ofexperience4OR4
- Master9s degree and 4 years of experience4OR4
- Bachelor9s degree and 6 years of experience4OR4
- Associate9s degree and 10 years of experience4OR4
- High school diploma / GED and 12 years of experience
Preferred Qualifications
- Deepexpertisein machine learning, deep learning, and Generative AI (LLMs, transformers, embeddings, fine-tuning techniques).
- Proventrack recordofleading and delivering production-grade ML/GenAI systems end-to-endwith measurable business impactwith strong experience in designingscalable system architecturesfor ML and GenAI, including distributed systems and high-throughput pipelines.
- ExpertiseinMLOps/LLMOpsecosystems(MLflow, Kubeflow, Airflow, CI/CD, Docker, Kubernetes).
- Strong systemdesign, architecture, and problem-solving skills with the ability tooperateindependently and lead large initiatives.
- Demonstratedproficiencyinleveragingcloud platforms (AWS, Azure, GCP) for data engineering solutions. Strong understanding of cloud architecture principles and cost optimization strategies.
- Provenability to mentor and guide junior and mid-level engineers (L4/L5).
Good-to-Have Skills:
- Cloud certifications (AWS, Azure, or GCP) are a plus
- Strong experience with big data ecosystems, including Apache Spark, Hadoop, and large-scale distributed data processing
- Deep expertise in data engineering, including building and optimizing scalable data pipelines and platforms using Databricks, Spark, SQL, and Python
- Advanced proficiency in Python and modern ML/AI frameworks (e.g., PyTorch, TensorFlow, Hugging Face, LangChain, or similar)
- Experience designing robust evaluation and validation frameworks, including automated evaluations, human-in-the-loop systems, safety testing, and monitoring
- Extensive experience with Retrieval-Augmented Generation (RAG) architectures, vector databases, and knowledge-grounded AI systems
- Strong understanding of agentic AI frameworks, including orchestration, planning, memory management, and tool integration
- Solid foundation in statistical modeling, experimentation design (A/B testing), and causal inference
- Experience with NLP, semantic search, embeddings, and vector search systems
- Familiarity with Responsible AI practices, including fairness, explainability, governance, and regulatory compliance
- Hands-on experience with cloud-native AI/ML services across AWS, Azure, or GCP, including cost and performance optimization
- Experience with the Databricks Lakehouse platform for enterprise-scale data engineering, ML, and GenAI workloads
- Exposure to advanced evaluation techniques such as red-teaming, adversarial testing, and synthetic data generation
- Strong experience in data modeling and performance tuning for both OLAP and OLTP systems
- Hands-on experience with workflow orchestration tools such as Apache Airflow, and distributed processing frameworks like Apache Spark
What you can expect of usAs we work to develop treatments that take care of others, we also work to care for your professional and personal growth and well-being. From our competitive benefits to our collaborative culture, we9ll support your journey every step of the way.
The expected annual salary range for this role in the U.S. (excluding Puerto Rico) is posted. Actual salary will vary based on several factors including but not limited to, relevant skills, experience, and qualifications.
In addition to the base salary, Amgen offers a Total Rewards Plan, based on eligibility, comprising of health and welfare plans for staff and eligible dependents, financial plans with opportunities to save towards retirement or other goals, work/life balance, and career development opportunities that may include:
- A comprehensive employee benefits package, including a Retirement and Savings Plan with generous company contributions, group medical, dental and vision coverage, life and disability insurance, and flexible spending accounts
- A discretionary annual bonus program, or for field sales representatives, a sales-based incentive plan
- Stock-based long-term incentives
- Award-winning time-off plans
- Flexible work models where possible. Refer to the Work Location Type in the job posting to see if this applies.
Apply now and make a lasting impact with the Amgen team.
careers.amgen.comIn any materials you submit, you may redact or remove age-identifying information such as age, date of birth, or dates of school attendance or graduation. You will not be penalized for redacting or removing this information.
Application deadline
Amgen does not have an application deadline for this position; we will continue accepting applications until we receive a sufficient number or select a candidate for the position.
Sponsorship
Sponsorship for this role is not guaranteed.