Job Description
The Command Center is the central monitoring Team responsible for 24x7 infrastructure operations, ensuring availability, performance, and stability of enterprise IT environments through proactive monitoring, incident management, and coordination across L2/L3 teams.
The team operates in a shift-based model (24x7 including weekends and holidays) to support business-critical environments with defined SLAs.
͏
Responsibilities:
A. Infrastructure Monitoring & Event Management
Perform real-time monitoring of infrastructure components (Servers, Network, Storage, Databases, Middleware).
Track alerts, thresholds, and system health across tools (e.g., monitoring/event platforms).
Identify potential failures and perform proactive issue detection and event correlation.
Example aligned activities:
Monitor CPU, Memory, Disk, uptime,
System and network performance
Monitor scheduled jobs, application availability, and backups.
B. Incident Management (L1 Command Center)
Act as the first line of support for all infrastructure-related incidents.
Log, categorize, prioritize, and assign tickets in ITSM tool.
Perform initial troubleshooting using SOPs and knowledge base.
Ensure adherence to SLA response and resolution timelines.
Escalate unresolved issues to L2/L3 teams based on defined escalation matrix.
Support Major Incident (P1/P2) bridge calls and coordination.
C. Command Center Operations Governance
Operate centralized Command Center model with cross-tower coordination (Compute, Network, Storage, DB, Middleware).
Provide single visibility across IT infrastructure.
Support continuous improvement of monitoring coverage and alert quality.
Ensure compliance with ITIL processes (Incident, Problem, Change).
D. Communication & Stakeholder Coordination
Act as Single Point of Contact (SPOC) during incidents and outages.
Provide timely updates to stakeholders and service owners.
Coordinate with:
• L2 / L3 technical teams
• Application teams
• Vendors/OEMs
Manage escalations and maintain communication cadence.
E. Reporting & Documentation
Maintain logs of alerts, incidents, and escalations.
Generate:
• Daily / Shift reports
• Incident summary reports
• SLA compliance reports
Update Knowledge Base (KB), Known Error Database (KEDB), and SOPs.
F. Preventive & Proactive Operations
Ensure execution of preventive maintenance tasks (as per SOP).
Identify recurring issues and trigger Problem Management.
Support backup verification, patch monitoring, and capacity alerts.
͏
Core Skills
1. Basic knowledge across Infra domains:
• Windows / Linux OS
• Networking fundamentals
• Storage & Backup
• Databases (basic understanding)
2. Understanding of Monitoring & ITSM tools
3. Knowledge of ITIL processes (Incident, Problem, Change)
Good to Have
1. Exposure to tools like OpsRamp, Solarwinds etc.
2. ITIL certification
͏
͏
Experience: 3-5 Years .
Reinvent your world. We are building a modern Wipro. We are an end-to-end digital transformation partner with the boldest ambitions. To realize them, we need people inspired by reinvention. Of yourself, your career, and your skills. We want to see the constant evolution of our business and our industry. It has always been in our DNA - as the world around us changes, so do we. Join a business powered by purpose and a place that empowers you to design your own reinvention.