Medallia

Senior Platform Software Engineer (Cache, Query and Compute Platforms)

Medallia$138K — $205K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in software, systems, platform, DevOps, or reliability engineering.
  • 3+ years operating distributed platforms in production environments.
  • Hands-on experience managing at least two of Redis, Kvrocks, Trino, or Spark.
  • Proven experience in incident management, root-cause analysis, and performance improvements.
  • Knowledge of high availability and fault tolerance in high-throughput systems.
  • Proficiency in Java, Go, Python, or similar programming language.
  • Experience in technical collaboration and architecture review processes.

Responsibilities

  • Operate and maintain Redis, Kvrocks, Trino, and Spark platforms.
  • Manage Redis Cluster and its operational tasks including failover and performance optimization.
  • Validate Kvrocks for topology and operational readiness.
  • Troubleshoot and optimize Trino clusters including query executions and integrations.
  • Manage and resolve issues in Spark clusters regarding resource usage.
  • Lead incident response, performing root-cause analysis, and tuning performance.
  • Develop operational automation, monitoring, and troubleshooting documentation.

Benefits

  • Competitive health and wellness benefits including medical, dental, and vision coverage.
  • 401(k) with company contributions.
  • Short-term and long-term disability insurance.
  • Life and AD&D insurance.
  • Paid parental leave and paid holidays.
Full Job Description
The Role and Team

We are seeking a hands-on Senior Platform Software Engineer with experience operating cache, distributed query, and compute platforms at scale. As a key member of the Platform Services team, you will help ensure the availability, reliability, performance, and operational readiness of Redis, Kvrocks, Trino, and Spark.

Responsibilities

  • Operate and maintain Redis, Kvrocks, Trino, and Spark platforms.
  • Manage Redis Cluster and Sentinel deployments, replication, failover, persistence, upgrades, and performance.
  • Validate Kvrocks topology, replication, failover, recovery, and operational readiness.
  • Troubleshoot Trino coordinators, workers, catalogs, query execution, S3 access, and Hive Metastore integrations.
  • Operate Spark clusters and troubleshoot drivers, executors, scheduling, failover, resource usage, and application health.
  • Lead production incident triage, root-cause analysis, and performance tuning.
  • Plan and validate upgrades, capacity, failover scenarios, and recovery procedures.
  • Review application architectures and production-readiness requirements.
  • Build operational automation, monitoring, alerts, documentation, and troubleshooting playbooks.
  • Participate in a periodic on-call rotation to maintain 24/7 reliability and performance of production services.

Candidates based in the Tysons vicinity will be prioritized as this role is Hybrid, 3 days per week onsite.

Qualifications

Minimum Qualifications
  • 5 years of experience in software, systems, platform, DevOps, or reliability engineering.
  • 3 or more years of experience operating distributed platforms in production.
  • Cache & Query Platforms: Hands-on operational experience managing and scaling at least two of the following platforms in production: Redis, Kvrocks, Trino, or Spark.
  • Experience leading production incident triage, root-cause analysis (RCA), and performance optimization for memory fragmentation, query execution bottlenecks, cluster failovers, and CPU/memory resource utilization.
  • High Availability & Fault Tolerance: Experience designing, testing and executing disaster recovery plans, automated node failovers, cluster re-sharding, and zero-downtime upgrades for high-throughput platforms
  • Demonstrated experience in Java, Go, Python, or a similar programming language.
  • Technical Collaboration & Review: Experience conducting formal architecture reviews, authoring technical design documents (e.g., RFCs/ADRs), and establishing operational runbooks across engineering teams.

Preferred Qualifications
  • Experience with additional platforms among Redis, Kvrocks, Trino, and Spark.
  • Experience operating distributed platforms on Kubernetes or cloud infrastructure.
  • Experience integrating Trino with object storage, Hive Metastore, or external data catalogs.
  • Experience validating topology, failover, recovery, and application-health behavior.
  • Experience with platform automation, observability, upgrades, capacity planning, and failure-mode testing.
  • Experience supporting large-scale, business-critical cache, query, or compute platforms.


Medallia is committed to equal pay and transparency. The annual base salary range for this position is $138,000 - $205,000. Please note that the salary range information provided is a general guideline and combines all of the distinct labor markets within the US. It is uncommon for an individual to be hired at or near the top of the range for their role and compensation decisions are dependent on a variety of factors. Medallia considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, candidate's work location, education/training, key skills, internal peer equity, external market data, as well as, market and business considerations when making compensation decisions.

Medallia also offers competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays. Benefits and eligibility may vary by location and role.

At Medallia, we celebrate diversity and recognize the value it brings to our customers and employees.

About Medallia

Medallia is a software company that provides customer experience management solutions. The company was founded in 2001 by Borge Hald and Amy Pressman and is headquartered in San Francisco, California. Medallia's software allows businesses to collect and analyze customer feedback across multiple channels, including email, social media, and mobile. The company's clients include some of the world's largest brands, such as Hilton, Delta Air Lines, and Mercedes-Benz. Medallia went public in 2019 and is traded on the New York Stock Exchange under the ticker symbol MDLA.
Learn more about Medallia
Size
2,037 employees
Market Cap
$5.3 billion
Industry
Net Income
-$148.6 million
Founded
2001
Revenue
$477.2 million
NASDAQ

Similar Jobs

More Jobs at Medallia

More Information Technology Jobs

Find similar Senior Platform Software Engineer (Cache, Query and Compute Platforms) jobs: