Operations Analyst (SRE) Category: Business Consulting, Strategy and Digital Transformation
Main location: United States, Virginia, Fairfax
Alternate Location(s): United States, Louisiana, Lafayette
Position ID:J0726-1858
Employment Type: Full Time
Position Description: CGI's Advantage Cloud Operations is an SRE-driven operating model, anchored on an Operations Control Plane that unifies telemetry, event management, automation, and IT Service Management (ITSM). The Operations Analyst is a hands-on Site Reliability Engineering (SRE) practitioner responsible for monitoring, triaging, and resolving incidents using correlated telemetry and automated runbooks. This role contributes to reducing operational toil through automation and disciplined problem management while supporting the reliability of CGI's Advantage platform.
This position can be performed from any CGI U.S. CSG office, with a preference for Lafayette, LA.
Your future duties and responsibilities: Operations & Reliability
• Serve in a hands-on operations role (SRE track) within the Cloud Operations team supporting the CGI Advantage platform.
• Take first-line ownership of monitoring, event triage, incident handling, and service request execution.
• Work from the operator portal utilizing correlated dashboards, one-click runbooks, and standardized operational workflows.
• Contribute to automation, ticket hygiene, and knowledge base quality to reduce repeat work.
• Progress toward the Senior Operations Analyst role through the SRE realignment career path.
Monitoring & Event Triage
• Monitor environment health using standardized dashboards and SLO-aligned alerts.
• Triage events in IAP by applying deduplication and suppression context while routing incidents according to severity taxonomy.
• Utilize distributed traces and correlated logs and metrics to isolate faults before escalation.
• Meet First Touch Resolution and L1-to-L2 escalation reduction targets.
Incident & Service Request Execution
• Execute approved runbooks and guard-railed self-service actions, including diagnostics, restarts, and environment tasks.
• Maintain complete, well-classified tickets using guided intake fields and consistent operational taxonomy.
• Coordinate activities within automatically created Microsoft Teams incident channels while documenting timelines and handoffs.
• Draft clear customer-facing status updates using standardized communication templates.
Continuous Improvement
• Identify recurring incidents as problem management candidates and support Root Cause Analysis (RCA) efforts with evidence and operational data.
• Develop small automation scripts using Python, Bash, or PowerShell and recommend new runbook candidates based on recurring operational tasks.
• Report configuration drift and validate changes against approved baselines under change control.
• Maintain runbooks and knowledge articles while leveraging LLM-assisted triage tools to improve operational efficiency.
Required qualifications to be successful in this role:• 2-5 years of experience in cloud or application operations, Network Operations Center (NOC), or L1/L2 production support environments.
• Working knowledge of Microsoft Azure and Kubernetes (AKS), including: Pods ,Deployments ,Logging ,Scaling concepts
• Experience with monitoring and logging platforms, including: Grafana-style dashboards ,Log search and analytics ,Alert management
• Basic scripting experience using: Bash,Python ,PowerShell
• Knowledge of IT Service Management (ITSM), including: Incident Management ,Service Request Lifecycle ,SLA awareness ,Quality ticket documentation
• Linux fundamentals
• Networking fundamentals, including: DNS ,TLS ,Load balancing
• Experience using Git
• Strong written and verbal communication skills for incident updates and operational handoffs.
Preferred Experience
. OpenTelemetry, Jaeger tracing, or Prometheus/Loki exposure.
. Ansible playbooks, Terraform basics, or GitHub Actions pipelines.
. Event management/AIOps platforms and runbook automation tools.
. PostgreSQL basics and SQL for operational queries and diagnostics.
. Certifications: AZ 900/AZ 104, CKA (entry), ITIL 4 Foundation.
CGI expects to accept applications for this position through 9/30/2026.
Other Information:
CGI is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. To support the ability to reward for merit-based performance, CGI typically does not hire individuals at or near the top of the range for their role. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $89,600.00 - $156,700.00.
At CGI we call our professionals "CGI Partners" to reinforce that all who join our team are, as owners, empowered to participate in the challenges and rewards that come from building a world-class company. CGI's benefits are offered to eligible professionals on their first day of employment to include:
• Competitive base salaries.
• Eligibility to participate in an attractive Share Purchase Plan (SPP) in which the company matches dollar-for-dollar contributions made by eligible employees, up to a maximum, for their job category.
• 401(k) Plan and Profit Participation for eligible CGI Partners.
• Flexibility and paid accrued vacation leave, ranging from 10 to 15 days per year, based on the job level, years of relevant prior experience, and years of service.
• 8 paid holidays per year & 2 floating holidays per year.
• Wellness and Well-being programs.
• Comprehensive medical, dental, and vision benefits
• Back-up child care, Pet insurance, a Member Assistance Program, a 529 college savings program, a personal financial management tool, lifestyle management programs and more.
For more detailed information about our benefits offerings visit Benefits | CGI Careers
Please note that the benefits listed above are subject to change based on the specific terms and conditions of the contract being supported.
#LI-KB5
Skills: - Azure
- Incident Management
- IT Service Management
- Kubernetes
- Python
- SRE (Site Reliability Engr.)