Role summary
The Assistant Manager – Cloud Operations supports day-to-day cloud operations, reliability, governance, security, and optimization of Microsoft Azure and Google Cloud Platform environments. The role collaborates with infrastructure, security, application, service desk, vendors, and partners to ensure cloud services remain stable, secure, and aligned with business requirements. Supports Site Reliability Engineering practices including service availability, monitoring, incident response, and operational excellence.
Responsibilities
- Support daily cloud operations across Microsoft Azure and Google Cloud Platform
- Monitor cloud infrastructure health, availability, performance, capacity, and incidents
- Support SRE practices including reliability monitoring, availability improvement, and incident response
- Manage incident response, service requests, escalations, and root cause analysis
- Support implementation and maintenance of cloud governance, policies, standards, and compliance controls
- Review and optimize cloud cost, usage, capacity, and resource cleanup
- Coordinate with vendors, managed service providers, and internal teams on operational issues
- Support backup, disaster recovery, high availability, and business continuity activities
- Maintain cloud documentation, operational runbooks, architecture diagrams, and knowledge base
- Support cloud security operations including IAM reviews, network security, and vulnerability remediation
- Assist with automation and Infrastructure as Code adoption using Terraform
- Track operational KPIs including availability, incident resolution, and SLA compliance
- Participate in change management, release coordination, and operational readiness reviews
- Support audits, compliance reviews, and risk remediation for cloud infrastructure
Qualifications
- Minimum 5 years hands-on cloud operations or cloud engineering experience
- Hands-on knowledge of Microsoft Azure and/or Google Cloud Platform
- Understanding of cloud networking, identity, compute, storage, and monitoring
- Experience with incident, problem, change, and service request management
- Familiarity with SRE concepts including availability, reliability, and incident response
- Familiarity with ITIL processes and IT service management tools
- Strong understanding of cloud security, IAM, RBAC, and encryption
- Ability to work with vendors, managed service providers, and cross-functional teams
- Strong troubleshooting, communication, documentation, and stakeholder coordination skills
Experience
Minimum 5 years of hands-on experience in cloud operations, cloud engineering, infrastructure operations, SRE, or related technical roles
Skills
- Microsoft Azure
- Google Cloud Platform
- Cloud networking
- Identity and access management
- Compute services
- Storage services
- Monitoring and alerting
- Backup and disaster recovery
- Cloud security
- Encryption
- Incident management
- Terraform
- PowerShell
- Python