The application window is expected to close on:
Job posting may be removed earlier if the position is filled or if a sufficient number of applications are received.
Role SummarySplunk, a Cisco company, helps organizations build digital resilience across security and observability. The Structured Store team builds and maintains critical data-storage infrastructure and services that power core Splunk workflows and applications. We help shape the future of structured storage at Splunk through reusable architectures, service APIs, and design patterns that set standards for how these systems are scaled, optimized, operated, and extended.
Our engineers own the full service lifecycle, from design and implementation through deployment, observability, on-call response, and continuous improvement. The team's portfolio spans graph, key-value, and related structured-storage services running on premises and across AWS, Azure, and GCP.
In this role, you will remain hands-on while leading end-to-end projects and medium-sized features, coordinating delivery across product, SRE, security, support, and engineering partners, and contributing to the team's defined roadmap.
What You'll Get to Do - Design, implement, test, and operate well-scoped projects and small-to-medium features within the team's defined structured-storage roadmap.
- Develop versioned APIs and service components for graph, key-value, and related structured-storage services, covering authentication, routing, placement, policy, caching, and database lifecycle operations; contribute actively to design and code reviews.
- Own delivery from source to production by using CI/CD, secure container images, Kubernetes configuration, database migrations, rollout validation, and rollback plans.
- Improve reliability, scalability, security, performance, and cost efficiency using metrics, traces, logs, profiles, load tests, and customer or product-success signals.
- Troubleshoot complex production issues, periodically participate in the team's on-call rotation, and when appropriate lead postmortems, root-cause analysis, and corrective actions.
- Coordinate project specifications and timelines, keep stakeholders informed, surface risks early, and develop practical remediation options when delivery changes.
- Share knowledge through documentation, reviews, and technical discussions; mentor peers or interns and help strengthen engineering practices across the team.
Minimum Qualifications - Bachelor's degree and 5+ years of related experience OR master's degree and 3+ years of related experience OR PhD.
- Experience developing production backend services and versioned APIs in Go, C++, or another systems programming language.
- Experience building or operating distributed services in production, including failure handling, scalability, or recovery.
- Experience with at least one production graph, relational, or NoSQL database, including data modeling, queries, schema or data migrations, transactions, or performance tuning.
- Experience deploying containerized services through CI/CD to Kubernetes or a comparable orchestration platform.
- Experience diagnosing production issues with logs, metrics, or traces and participating in on-call or incident response.
Nice-to-Have Qualifications If you meet the core qualifications but not every preferred item, we still encourage you to apply.
- Neo4j and Cypher experience, including graph modeling, clustering, topology, Bolt or Query API usage, database provisioning, page cache, and transaction metrics.
- PostgreSQL, Amazon Aurora, or RDS experience involving connection security, IAM authentication, migrations, tuning, backup, and restore.
- Experience with Helm, Kubernetes controllers or Operators, GitLab CI/CD, Terraform, Puppet, or similar delivery and configuration tooling.
- Experience operating services across one or more public clouds and on-premises environments, including service mesh, network policy, workload identity, or Vault integrations.
- Knowledge of database internals, storage engines, query planning, indexes, transaction logs, concurrency control, or large-scale performance testing.