Your role and responsibilities
As an Associate Data Engineer, you will help design, build, and improve data platforms, data products, and services that support analytics, machine learning, generative AI, and agentic AI solutions. You will work across data gathering, ingestion, transformation, storage, batch and real-time processing, semantic enrichment, retrieval, APIs, visualization, data quality, observability, and governance.
Collaborating closely with diverse teams, you will play an important role in selecting suitable data management systems and identifying the critical data needed for insightful analysis. As a Data Engineer, you will help tackle challenges related to database integration, data quality, and complex structured and unstructured datasets.
Key Responsibilities may include:
- Assisting designing and implementing scalable data architecture, data products, and management systems for modern cloud environments, analytics, and AI-enabled use cases.
- Work on optimizing existing data pipelines, retrieval indexes, and data services for improved performance, reliability, data quality, and freshness of AI-ready context.
- Collect, prepare, and analyze structured, semi-structured, and unstructured data to identify trends, providing clients with actionable insights that enhance marketing, operational, and business practices.
- Participate in troubleshooting data-related issues, working to solve data quality challenges, retrieval-quality gaps, processing failures, and inconsistencies affecting analytics, generative AI, or agentic workflows.
- Create visually compelling and user-friendly dashboards, reports, and observability views to communicate findings, pipeline health, and AI system insights to both technical and non-technical stakeholders.
- Ensure data integrity, accuracy, reliability, lineage, and access control through rigorous data cleaning, validation, preprocessing, cataloging, and governance practices.
- Work with project teams to prioritize and translate client requirements into current and future operational scenarios, processes, models, use cases, data products, APIs, tool interfaces, plans, and solutions; collaborate with clients, architects, and AI engineers.
- Present analytical findings, data quality insights, and recommendations clearly and concisely, demonstrating the value of data-driven and AI-enabled decision-making to clients.
- Work with cross-functional teams to tackle complex business problems, utilizing data expertise across cloud platforms, RAG, vector search, APIs, responsible AI, and agent orchestration patterns while staying current on modern data stack trends.
Consulting And Collaboration Skills
- Analyze business processes and application portfolios to identify where teams need better data, context, tools, automation, or process optimization.
- Use Agile ways of working, planning, and project-management practices to deliver production-ready data and AI solutions.
- Communicate with curiosity, clarity, and empathy; ask insightful questions about data meaning, risk, governance, and business impact.
Required education
High School Diploma/GED
Preferred education
Bachelor's Degree
Required technical and professional expertise
- Familiarity with one or more programming or query languages such as Python, SQL, Java, Scala, or JavaScript.
- Foundational understanding of data engineering concepts, including data pipelines, databases, APIs, distributed processing, ETL/ELT, data modeling, or data products.
- Basic understanding of cloud computing environments such as AWS, Azure, Google Cloud, IBM Cloud, or similar platforms.
- Ability to apply foundational statistical, machine learning, or information retrieval concepts to data preparation, analysis, search, or model-support work.
- Interest in AI/ML, generative AI, agentic AI, intelligent automation, RAG, LLM-powered systems, or enterprise AI platforms.
- Strong analytical thinking, problem solving, adaptability, teamwork, and communication skills.
- Willingness to travel up to 100%, based on project requirements.
Preferred technical and professional experience
- Bachelor's degree in a related field such as Computer Science, Data Science, Statistics, Mathematics, MIS, Engineering, AI/ML, or another quantitative field.
- Coursework, projects, internship experience, or portfolio work involving data engineering, software engineering, cloud platforms, analytics, AI/ML, or modern data products.
- Experience with technologies such as Spark, Hadoop, Kafka, Airflow, dbt, Databricks, Snowflake, Delta Lake, Linux, or similar platforms.
- Familiarity with LLMs, embeddings, vector databases such as Pinecone, Weaviate, orpgvector, retrieval systems, prompt workflows, RAG, and AI agent architectures.
- Exposure to orchestration and agentic AI frameworks such asLangChain,LangGraph,LlamaIndex, Semantic Kernel, AutoGen, MCP-based tooling, or similar tools.
- Understanding of Git, containers, APIs, testing frameworks, CI/CD, Kubernetes, observability, data governance, privacy, security, or responsible AI practices.
- Preferred certifications, coursework, or credentials such asSnowProCore, Google Associate Cloud Engineer, Google Data Engineer coursework, or related cloud/data credentials.
- Familiarity with AI coding tools and GenAI concepts, including prompt engineering, RAG, fine-tuning, and model evaluation.
OTHER RELEVANT JOB DETAILSIBM offers a competitive and comprehensive benefits program. Eligible employees may have access to:
- Healthcare benefits including medical & prescription drug coverage, dental, vision, and mental health & well being
- Financial programs such as 401(k), cash balance pension plan, the IBM Employee Stock Purchase Plan, financial counseling, life insurance, short & long- term disability coverage, and opportunities for performance based salary incentive programs
- Generous paid time off including 12 holidays, minimum 56 hours sick time, 120 hours vacation, 12 weeks parental bonding leave in accordance with IBM Policy, and other Paid Care Leave programs. IBM also offers paid family leave benefits to eligible employees where required by applicable law
- Training and educational resources on our personalized, AI-driven learning platform where IBMers can grow skills and obtain industry-recognized certifications to achieve their career goals
- Diverse and inclusive employee resource groups, giving & volunteer opportunities, and discounts on retail products, services & experiences
We consider qualified applicants with criminal histories, consistent with applicable law.
This position was posted on the date cited in the key job details section and is anticipated to remain posted for 21 days from this date or less if not needed to fill the role.
IBM will not be providing visa sponsorship for this position now or in the future. Therefore, in order to be considered for this position, you must have the ability to work without a need for current or future visa sponsorship.
The compensation range and benefits for this position are based on a full-time schedule for a full calendar year. The salary will vary depending on your job-related skills, experience and location. Pay increment and frequency of pay will be in accordance with employment classification and applicable laws. For part time roles, your compensation and benefits will be adjusted to reflect your hours. Benefits may be pro-rated for those who start working during the calendar year.