Full Job Description
As a DBA Engineer, you will assist customers in developing, deploying, operating, and continuously improving cost-effective solutions on OSS and BSS Telecom Applications. The ideal candidate will be responsible for the performance, integrity, and security of our databases. This role involves installation, configuration, maintenance, performance monitoring, security, backup and recovery, and troubleshooting.
Responsibilities
• Should have basic understanding of Telecom OSS domain.
• Should have understanding on ITIL Concepts.
• Experience working in a 24x7 production support model with excellent troubleshooting skills. • Should have understanding on Ticket prioritizing based on impact/urgency, SLA, RCA
• Should be ready to work in 24/7 rotational shifts with flexibility timings.
• Good communication skill.
• Experience with scripting language is preferred. • Respond to reported service incidents/requests and initiate the incident management process
Keep users informed about their incidents' status at agreed intervals • Verify resolution with users and resolve incidents in tool
• Log all incidents/service requests and their resolution to identify recurring issues • Perform daily triage of incidents
• Configure threshold value alerting, triggers, and remediation
• Monitor network utilization and system health checks
• Maintain and update reports and correspondence related to the work • Prioritize incidents according to their urgency and impact on the business
• Analyse and escalate the incident to the next level (L2) if that incident is not in the scope of L1 or L1.5
Requirements
• Minimum of 10+ years of Oracle, SQL database management systems in testing, implementation, maintenance and administration in a multiple platform environment.
• A Bachelor's or Master's degree in Computer Science or Engineering or AI Engineering, or a related field is typically required.
• Upgrades of Oracle, SQL , PostgreSQL database migrations and updates and application of database patches.
• Backup and restores, export and import of data to and from Oracle, SQL run time configuration of Oracle, SQL.
• Develop and implement backup and recovery strategies to protect against data loss.
• Monitoring at server, database, collection level, and using various monitoring tools related to Oracle, PostgreSQL.
• Strong understanding of SQL architecture and internals.
• Proficiency in SQL and database/data design and architectural principles and methodologies.
• Familiarity with Linux/Unix operating systems.
• Scripting skills (e.g., Bash, Python) for automation.
• Supports multiple services and multiple databases of medium complexity with multiple concurrent users, ensuring control, integrity and accessibility of data.
• Implements multiple projects along with the production support.
• Will perform on-call support for critical needs.
• Experience with database monitoring tools.
• Implement performance tuning strategies to ensure optimal database response times.
• Responsible for providing support to customers by researching, diagnosing, troubleshooting issues, and resolving incidents and providing support for software bugs and other technical problems.
• Must have experience in the role of 24/7 Production Support and Maintenance activities.
• Design, build, Manage and maintain CI/CD tools to accelerate software development and deployment.
• Identify, troubleshoot, and resolve infrastructure issues in development, testing, and production environments.
• Perform system tests for security, performance, and availability.
• Automate build, test, and deployment processes.
• Automate repetitive tasks using scripting languages like Python, Bash, or PowerShell.
• Design, implement, and manage cloud infrastructure on platforms like AWS, Azure, or GCP.
• Provision and configure servers, databases, and other infrastructure components.
• Implement infrastructure-as-code (IaC) using tools like Terraform or CloudFormation.
• Implement monitoring and logging systems to track application performance and identify potential issues.
• Develop alerting mechanisms to notify teams of critical issues.
• Ticket management efficiencies i.e. response time, resolution time, providing regular updates, SLAs.
• Escalate incidents at risk of breaching SLAs to the responsible teams
• Scale up to demonstrate flexibility in handling business criticality and handling spikes
• Propose automation opportunities for repetitive tasks to improve change management process
• Conduct Root Cause Analysis (RCA) post restoration of service
• Coordinate with vendors to resolve hardware and software incidents and follow-up until service is restored and ticket closure
• Learn and comply with validation requirements, standard-operating procedures (SOPs), project quality model (PQM), and change control maintenance for product life cycle
• Build relationships with stakeholders and understand how to meet their IT requirements while adhering to Prodapt's best practices