Job Description
We are looking for a Infrastructure Operations SRE to own Infra infrastructure operations and stability platform engineering. The role covers day-to-day operations of Kubernetes clusters, databases and middleware, and application releases and changes, as well as data onboarding and operations for Infra stability platforms, ensuring business platforms run stably in an independent cloud environment.
Responsibilities
- Perform daily monitoring, health checks, and capacity management of Kubernetes clusters; track node CPU/memory/disk utilization, Pod status, and database/middleware metrics; handle alerts, collect logs to identify risks, and execute scaling or disk cleanup as needed;
- Respond to monitoring alerts and business feedback; troubleshoot failures of K8s nodes, middleware components, and application instances; perform fault isolation, hardware decommission and repair, workload migration, and capacity expansion to ensure fast recovery;
- Maintain databases and middleware including MySQL/MongoDB/PostgreSQL/Kafka/RocketMQ/Redis/Nacos; perform backup/restore and disaster recovery drills; manage gateway configuration including domains, routing rules, and certificates;
- Own application initialization, version releases and configuration changes, elastic scaling for traffic fluctuations, and application migration during node failures; maintain a high change success rate;
- Manage permission provisioning for middleware and databases, servers/network/Infra consoles, and account synchronization to keep daily operations requests handled efficiently;
- Drive production safety and security: vulnerability remediation and hardening, incident coordination, severity assessment and postmortems, emergency response, and day-to-day operations support;
- Own data onboarding and operations for Infra stability platforms, covering device management, energy monitoring, emergency drills, and key process control; support platform iteration and ensure data completeness and accuracy;
- Ensure the stability of operations platforms deployed in an independent cloud environment; build platform monitoring, emergency playbooks, and incident response mechanisms; drive high-availability architecture and disaster recovery capability.
Position Requirement
Minimum qualifications:
Bachelor's degree or above in Computer Science or a related field, with 3+ years of experience in infrastructure operations, SRE, or Infra operations;
Solid Linux administration and Shell/Python scripting skills; familiar with Kubernetes orchestration and common cloud-native monitoring; hands-on experience with cluster health checks, troubleshooting, and capacity management;
Proficient in at least one mainstream database or middleware such as MySQL, Redis, or Kafka, with an understanding of high-availability architecture, backup/restore, and common performance issue handling;
Solid development skills in at least one of Python/Go/Java, able to independently build operations automation tools and platform features;
Familiar with application release/change processes and gateway configuration management, with disciplined change and rollback practices;
Strong security awareness; familiar with permission management, vulnerability remediation, and incident response fundamentals;
Strong sense of ownership, excellent cross-team collaboration and communication skills, and the ability to respond effectively to emergencies.
Preferred qualifications:
Experience in Infra / Infra infrastructure operations;
Experience building or developing stability platforms or operations platforms;
Experience operating large-scale multi-cluster or multi-site environments.
The pay range for this position at commencement of employment is expected to be between $145,200/year and $238,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.
If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.