“I can succeed as a Senior Platform Engineer – Observability at Capital Group.”
As a Senior Platform Engineer on our observability platform team, you’ll design, build, and operate the observability capabilities that thousands of engineers rely on to understand the health, performance, and behavior of their applications and infrastructure. You’ll build on open telemetry standards to give teams a unified, vendor-flexible experience across metrics, events, logs, and traces—turning raw signals into actionable insight that accelerates troubleshooting, strengthens reliability, and improves the developer experience.
This is a hands-on senior engineering role. You’ll engineer telemetry collection and pipelines, automate onboarding so teams get observability out of the box, and deliver Observability-as-Code—monitors, dashboards, and alerts managed as versioned, reusable modules. A central goal is a true single pane of glass: you’ll correlate metrics, events, logs, and traces into one unified experience so engineers can move seamlessly from signal to root cause without switching tools. You’ll integrate with leading observability backends (for example, a SaaS APM/metrics platform such as Datadog alongside log platforms and open-source stacks like Prometheus and Grafana) while keeping the architecture standards-based and portable, so we’re never locked to a single vendor. You’ll define golden-signal and SLI/SLO standards, tune cost and cardinality, advance AIOps and anomaly detection, and coach engineering teams to raise observability maturity across the organization. You’ll also develop with AI—using AI-assisted coding tools and agentic workflows to build and refactor platform tooling faster, while keeping quality and security high.
“I am the person Capital Group is looking for.”
- You have a bachelor’s degree in Computer Science, Engineering, Information Systems, or equivalent technical experience
- You have at least 8 years in Software Engineering, Platform Engineering, SRE, or Cloud Engineering, with a strong software background designing, building, and operating production-grade systems at scale
- You have strong software engineering fundamentals, along with a solid foundation in networking, security, and cloud-native architectures, and can apply them to solve complex platform engineering challenges at scale.
- You understand the trade-offs between reliability, performance, and cost, and apply sound engineering judgment when designing platform solutions.
- You have deep, hands-on observability platform experience—instrumenting applications and building telemetry pipelines and backends across metrics, events, logs, and traces (MELT), grounded in OpenTelemetry standards with a vendor-neutral, portable design mindset—using modern tooling such as a leading SaaS APM/metrics platform (e.g., Datadog), open-source stacks (Prometheus, Grafana, OpenTelemetry Collector), and log platforms (e.g., Splunk, CloudWatch)
- You have proven experience delivering unified, single-pane-of-glass observability—correlating metrics, events, logs, and traces (including trace-to-log correlation and consistent tagging/context propagation) so engineers move from signal to root cause in one experience without switching tools
- You are an expert in at least one of Python or Go (our primary languages for platform tooling, collectors, and automation) and are comfortable reading and instrumenting services in additional languages such as Java, .NET (C#), and Node.js/TypeScript; strong telemetry query-language skills (e.g., PromQL, SQL, or equivalent) are expected
- You develop with AI as part of your craft—using AI-assisted coding tools (e.g., GitHub Copilot or equivalent) and agentic workflows to accelerate development, testing, and refactoring of platform tooling—while applying sound engineering judgment to review, validate, and secure AI-generated code in line with approved enterprise AI platforms and controls
- You have strong Infrastructure-as-Code skills (Terraform/OpenTofu) and deliver Observability-as-Code—monitors, dashboards, and alerts as versioned, reusable modules deployed through CI/CD—plus telemetry collection and pipeline engineering (collector/agent fleet management, data hygiene and tagging standards, and cost/cardinality/sampling optimization) on cloud-native AWS (EKS/Kubernetes, containers)
- You define reliability and observability standards—golden signals, distributed tracing/context propagation, SLIs/SLOs, structured logging, RUM/synthetics—and apply AIOps/anomaly detection to accelerate detection and root-cause analysis, while leading technical workstreams independently, mentoring engineers, and communicating clearly with technical teams and stakeholders (Agile/SCRUM)
“I can apply in less than 4 minutes.”
You’ve reviewed this job posting and you’re ready to start the candidate journey with us. Apply now to move to the next step in our recruiting process. If this role isn’t what you’re looking for, check out our other opportunities and join our talent community.
“I can learn more about Capital Group.”
Charlotte Base Salary Range: $136,749-$218,798
In addition to a highly competitive base salary, per plan guidelines, restrictions and vesting requirements, you also will be eligible for an individual annual performance bonus, plus Capital’s annual profitability bonus plus a retirement plan where Capital contributes 15% of your eligible earnings.
You can learn more about our compensation and benefits here.