Job Summary
We are looking for a Senior Application Operations & Performance Architect to provide technical leadership for a portfolio of approximately 10 enterprise applications. This role will serve as the technical bridge between Application Architecture, Development, DevOps, Production Operations, and Performance Engineering teams. The individual will be responsible for application reliability, production troubleshooting, performance optimization, observability, and identifying technical improvements across the application portfolio. This is a hands-on technical leadership role requiring strong application engineering knowledge combined with production operations, DevOps, cloud, and performance engineering experience.
Key Responsibilities
• Provide technical ownership and operational leadership across a portfolio of approximately 10 enterprise applications.
• Work closely with Application Architects and Development teams to understand application architecture, dependencies, technical constraints, and opportunities for improvement.
• Review application architecture and production behavior and recommend solutions to improve performance, scalability, reliability, resiliency, and operational efficiency.
• Use Dynatrace/APM, logs, metrics, traces, SQL analysis, and other diagnostic tools to proactively identify application and infrastructure bottlenecks.
• Troubleshoot complex production issues across Java applications, APIs, microservices, databases, Kubernetes, AWS, and supporting infrastructure.
• Lead technical root cause analysis for recurring or high-impact production and performance issues and drive corrective and preventive actions.
• Analyze application performance, including response time, latency, throughput, resource utilization, database performance, JVM behavior, and service dependencies.
• Identify recurring operational issues and work with Development, Architecture, and DevOps teams to engineer permanent solutions rather than relying on manual operational workarounds.
• Partner with DevOps teams to improve CI/CD, deployment reliability, automation, monitoring, alerting, and production readiness.
• Define and improve application observability, dashboards, alerts, and operational health metrics.
• Support capacity planning and proactively identify potential scalability and performance risks.
• Provide technical direction and guidance to the offshore Operations and Performance Engineering team.
• Drive continuous improvement initiatives to reduce incidents, recurring defects, operational effort, and Mean Time to Resolution (MTTR).
• Act as the senior technical point of contact for application operations and performance discussions with client architects and engineering leadership.
Required Qualifications
• 10+ years of overall IT experience with a strong application engineering and production operations background.
• Strong hands-on experience with Java/J2EE.
• Strong understanding of microservices, APIs, distributed systems, and enterprise application architecture.
• Hands-on experience with SQL, query analysis, and database performance optimization.
• Strong experience with Dynatrace, New Relic, AppDynamics, or similar APM/observability platforms.
• Strong production troubleshooting, incident management, and Root Cause Analysis (RCA) experience.
• Understanding of DevOps, CI/CD, deployment, and release processes.
• Experience troubleshooting application performance across application, API, database, infrastructure, and cloud layers.
• Ability to analyze logs, metrics, traces, application behavior, and system dependencies.
• Strong communication skills with the ability to work directly with application architects, engineering leaders, developers, operations teams, and business stakeholders.
• Experience working in an onsite-offshore delivery model.
Preferred Qualifications
• Experience with AWS.
• Experience with Kubernetes/EKS.
• Experience with AWS Lambda.
• Experience with Docker and containerized applications.
• Experience with Jenkins, GitLab, GitHub Actions, or similar CI/CD technologies.
• Experience with Splunk, CloudWatch, ELK, or similar logging platforms.
• Experience with JVM profiling and tuning.
• Experience with infrastructure and cloud performance optimization.
• Experience with capacity planning and scalability.
• Experience with JMeter, Gatling, LoadRunner, or similar performance engineering tools.
• Experience with SRE/reliability engineering practices.
• Experience with Infrastructure as Code and automation.