Database Reliability Engineer & Administrator

TCN

$90K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT, or related field.
  • 3+ years as a DBA, DBRE, or SRE in a Linux environment.
  • Advanced knowledge of PostgreSQL internals and replication.
  • Experience managing ClickHouse and OLAP workloads.
  • Expertise in Google Cloud Platform (GCP) and Kubernetes for stateful workloads.
  • Proficient in scripting (Bash, Python), with familiarity in Go a plus.
  • Strong communication skills to convey technical details to both developers and stakeholders.

Responsibilities

  • Collaborate with developers on schema migrations and database deployment for high availability.
  • Proactively monitor and optimize performance of PostgreSQL and ClickHouse.
  • Automate database provisioning and configuration with tools like Terraform and Ansible.
  • Implement high-availability solutions and ensure effective backup/recovery strategies.
  • Lead root-cause analysis for database incidents and troubleshoot complex issues.
  • Maintain observability solutions including dashboards and alerts for database health.
  • Participate in 24/7 on-call rotation for database incident response.

Benefits

  • Medical, Dental, and Vision Insurance
  • Life Insurance
  • 401k with employer match
  • Paid time off and 11 scheduled paid holidays
  • Weekly lunches and free drinks/snacks
  • Casual dress code and flexible work environment
Full Job Description
Database Reliability Engineer & Administrator (PostgreSQL & ClickHouse)

TCN is looking for a Database Reliability Engineer (DBRE) to join our team in Saint George, Utah. In this role, you will be the guardian of our data layer, ensuring that TCN's global production databases are performant, scalable, and resilient.

While you will share the core DNA of a Site Reliability Engineer, your specific mission is to optimize and manage our high-traffic PostgreSQL environments and our high-performance ClickHouse analytical clusters. You will bridge the gap between application development and database operations, ensuring our data infrastructure keeps pace with our global growth.
Key Responsibilities
  • Database Architecture & Deployment: Collaborate with developers to design schema migrations and deploy database changes that maintain high availability. Assist in the architectural design of PostgreSQL and ClickHouse clusters to ensure they meet scaling requirements.
  • Performance Tuning & Optimization: Proactively monitor and tune PostgreSQL (query optimization, indexing strategies, vacuuming) and ClickHouse (merge tree optimizations, shard/replica management) to ensure sub-second latency for our clients.
  • Infrastructure as Code: Automate the provisioning and configuration of database clusters using tools like Terraform, Ansible, or Kubernetes Operators.
  • Resilience & Failure Management: Manage high-availability (HA) solutions (e.g., Patroni, PGBouncer) and ensure robust backup/recovery strategies are tested and functional.
  • Deep Troubleshooting: Lead root-cause analysis for complex database incidents. Debug locking issues, replication lag, and resource contention in a cloud-native environment.
  • Observability: Build and maintain dashboards and alerting for database health, focusing on SLIs/SLOs related to data consistency and query performance.
  • Incident Response: Participate in a 24/7 on-call rotation, serving as the subject matter expert for database-related outages.
Qualifications
  • Education: Bachelor's degree in Computer Science, Information Technology, or a related field.
  • Experience: 3+ years in a Linux environment as a Database Administrator (DBA), DBRE, or SRE with a heavy focus on data systems.

Database Expertise:
  • PostgreSQL: Advanced knowledge of internal mechanics, replication (physical/logical), and extension management.
  • ClickHouse: Experience managing OLAP workloads, understanding of MergeTree engines, and distributed table configurations.
  • Cloud & Containers: Demonstrated experience with Google Cloud Platform (GCP) and running stateful workloads inside Kubernetes.
  • Linux Mastery: Deep understanding of the Linux OS, specifically how kernel parameters, storage I/O, and networking impact database performance.
  • Automation: Proficient in scripting (Bash, Python) and configuration management. Familiarity with Go is a significant plus.
  • Networking: Knowledge of TCP/IP, TLS encryption for data in transit, and load balancing (L4/L7).
  • Soft Skills: Excellent communication skills; the ability to explain "why" a query is slow to a developer and "what" the business impact is to a stakeholder.

Our benefits include:
  • Medical Insurance (HDHP with HSA)
  • Dental Insurance
  • Vision Insurance
  • Life Insurance
  • 401k with employer match
  • Competitive salary
  • Paid time off
  • Paid holidays (11 scheduled)
  • Weekly lunches; free drinks and snacks
  • Casual dress and flexible work environment

Similar Jobs

More Jobs at TCN

More Information Technology Jobs

Find similar Database Reliability Engineer & Administrator jobs: