We are looking for a Senior Software Engineer who can design and build automation that streamlines cluster provisioning, configuration, and lifecycle management, extend observability so that on-call engineers and platform users can quickly understand cluster health and performance across a growing fleet, and build self-service tooling that reduces the need for manual, one-off engineering support during onboarding. You'll work across the CI/CD pipeline, cloud infrastructure, and Solr cluster management stack, and collaborate closely with both the platform team (managed service, deployment, and reliability) and the SRE team (alarming, capacity, on-boarding).
Essential functions
- Design and build automation for onboarding new Solr use cases and tenants onto the managed platform, reducing manual setup and configuration steps.
- Extend and improve observability for the Solr platform, including metrics, dashboards, alerting, and tracing, so cluster health and performance are visible at scale.
- Design and develop tooling that gives platform consumers self-service visibility into their clusters, capacity, and configuration.
- Collaborate with the Platform and SRE sub-teams to identify gaps in tooling, automation, and monitoring as the platform scales, and prioritize the highest impact improvements.
- Troubleshoot and resolve system issues, improving system availability and performance as the number of onboarded use cases grows.
- Contribute to runbooks, documentation, and internal tooling that make the platform easier for the broader team to operate and support.
Qualifications
- Strong proficiency in Java, including backend development.
- Experience building RESTful services and resilient, high-performance backend components.
- Experience designing and implementing CI/CD automation using Gradle, Jenkins, Spinnaker, and GitHub.
- Solid working knowledge of AWS services, including deploying and managing applications in the cloud.
- Proficiency with Kubernetes for container orchestration, deployment, and application monitoring.
- Experience building or operating observability tooling (metrics, dashboards, alerting, logging, or tracing) for distributed systems.
- Experience designing systems for high availability, scalability, and performance.
- Comfortable working independently on ambiguous problems and collaborating across multiple engineering teams.
- Bachelor's degree in Computer Science, a related technical field, or equivalent practical experience
Would be a plus
- Experience with Apache Solr, Elasticsearch, or other distributed search and indexing technologies.
- Experience integrating with Apache Kafka and other message queuing systems.
- Experience building internal developer platforms or self-service infrastructure tooling.
- Experience with PagerDuty or similar alerting and incident management tooling.
- Familiarity with large-scale multi-tenant infrastructure and the operational challenges of onboarding new consumers safely.
We offer
- Opportunity to work on cutting-edge projects
- Work with a highly motivated and dedicated team
- Competitive salary
- Flexible schedule
- Benefits package - medical insurance, vision, dental, etc.
- Corporate social events
- Professional development opportunities
- Well-equipped office
- Please note that all onboarding must occur in person and you may be asked to travel to attend