Platform Reliability Engineer

Appnovation Technologies

• $110K — $130K *
Miami, FL 33186In-Person
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years in SRE, platform reliability, or observability engineering with hands-on AWS expertise.
  • Proficient in OpenTelemetry (SDKs, collectors, exporters) and distributed tracing techniques.
  • Experience with Amazon CloudWatch and at least one enterprise observability platform (e.g., Datadog, Splunk, Grafana).
  • Demonstrated ability to define and implement SLOs, error budgets, and alerting strategies.
  • Skilled in creating AWS cost visibility and FinOps reporting using Cost Explorer, CUR, tagging strategies.
  • Proficient in scripting languages such as Python, Go, or TypeScript.
  • Strong communication skills, capable of conveying data-driven insights to technical and non-technical stakeholders.

Responsibilities

  • Build OpenTelemetry instrumentation standards to facilitate adoption by teams.
  • Integrate AgentCore Observability with existing tools like Datadog and CloudWatch.
  • Establish end-to-end tracing to pinpoint issues across agent interactions.
  • Define SLOs for key metrics and implement noise-reducing alerting mechanisms.
  • Monitor token usage and compute costs through dashboards and alert systems.
  • Develop detailed runbooks for incident response and conduct post-incident analysis for improvement.

Benefits

  • Opportunity to work in a collaborative and innovative environment.
  • Access to advanced tools and technologies in observability and reliability.
  • Engagement in a high-impact project for a leading global life sciences client.
  • Chance to shape and influence the reliability platform from inception.
  • Professional development opportunities in a consulting context.
Full Job Description
We're putting together a dedicated delivery pod to build and run an internal agent platform on AWS Bedrock AgentCore for a global life sciences client. The pod works as one team with the client's engineers to deliver the platform other teams will build their agents on. In this role you make the platform visible and dependable. You'll build the OpenTelemetry instrumentation, observability integrations, tracing, SLOs, alerting, and cost monitoring for token and compute spend. When an agent is slow, wrong or expensive, your work is how the team finds out and fixes it. This is a named-team engagement, so the person we propose is the person who starts. ROLE RESPONSIBILITIES • Instrumentation: Build OpenTelemetry instrumentation standards for agents, tools and platform services, and make it easy for teams to adopt. • Observability Integration: Connect AgentCore Observability and CloudWatch with the client's existing observability tools (e.g., Datadog, Splunk, Grafana). • Tracing: Set up end-to-end tracing across agent steps, model calls, tool calls and multi-agent handoffs so issues can be traced to their source. • SLOs and Alerting: Define SLOs for latency, availability and error rates, and build alerting that is useful and not noisy. • Cost Monitoring: Track token usage and compute spend by team, agent and environment, with dashboards, budgets and alerts for unexpected spikes. • Incident Readiness: Write runbooks, support incident response and run post-incident reviews that lead to real fixes. QUALIFICATIONS • 6+ years in SRE, platform reliability or observability engineering, with strong hands-on AWS experience. • Hands-on experience with OpenTelemetry (SDKs, collectors, exporters) and distributed tracing. • Experience with Amazon CloudWatch and at least one enterprise observability platform (Datadog, Splunk, Grafana, New Relic or similar). • Experience defining and running SLOs, error budgets and alerting strategies. • Experience building cost visibility and FinOps reporting on AWS (Cost Explorer, CUR, tagging strategies). • Scripting and coding skills in Python, Go or TypeScript. • Clear communication and able to turn data into decisions for technical and non-technical audiences. PREFERRED QUALIFICATIONS • Experience monitoring LLM or agent-based applications, including token usage, latency and quality signals. • Familiarity with LLM observability tools (e.g., Langfuse, Arize, LangSmith) or AgentCore Observability. • Experience with infrastructure as code (Terraform or AWS CDK) for monitoring resources. • Experience in pharma, life sciences or another regulated industry. WHO YOU ARE • You want to know what's happening in a system before a user tells you • You build alerts people trust and dashboards people actually open • You treat cost as a reliability concern, not an afterthought • You stay calm in incidents and focus on learning afterward • You work well inside a client team and build trust quickly • You have prior experience in consulting • Prior experience and connections in the Life Sciences industry is preferred

More Jobs at Appnovation Technologies

More Information Technology Jobs

Find similar Platform Reliability Engineer jobs: