About Comprinno
Comprinno is a NASSCOM-incubated company headquartered in Bangalore, with offices in Pune, Coimbatore, and the United States. We specialize in cloud transformation, DevOps, infrastructure automation, and managed cloud operations, enabling organizations to build scalable, secure, and high-performing cloud environments on AWS.
Our flagship SaaS platform, Tevico, is an intelligent cloud governance and observability platform that helps enterprises improve reliability, optimize costs, strengthen security, and automate operational workflows across AWS environments.
As an AWS Advanced Consulting Partner, Comprinno helps customers migrate, modernize, secure, and manage cloud environments while adopting emerging technologies such as Data Analytics, AI, and Generative AI.
About the Role
We are seeking a highly skilled CloudOps Engineer – L3 to lead complex cloud operations, reliability engineering, automation initiatives, and customer success activities within our Managed Services practice.
This role serves as the highest technical escalation point within CloudOps and is responsible for ensuring operational excellence, platform reliability, automation maturity, security compliance, and cloud optimization across customer environments supporting customers 24/7 in rotational shifts.
The ideal candidate combines deep AWS expertise, strong troubleshooting capabilities, automation experience, and customer-facing consulting skills. You will work closely with Cloud Engineering, Security, Presales, and Delivery teams to drive operational maturity and continuous improvement.
Key Responsibilities
Advanced Cloud Operations
- Own operational health, reliability, and availability of customer AWS environments.
- Manage complex production environments and mission-critical workloads.
- Lead troubleshooting of high-severity incidents and infrastructure outages.
- Serve as the final escalation point for L1 and L2 operational issues.
- Drive resolution of P1 and P2 incidents within SLA commitments.
- Coordinate with AWS Support and third-party vendors during critical escalations.
Reliability Engineering & Platform Stability
- Establish and improve Site Reliability Engineering (SRE) practices.
- Define and track:
- Improve platform resilience, scalability, and operational efficiency.
- Conduct failure analysis and proactively eliminate recurring operational issues.
- Drive operational readiness reviews and reliability assessments.
Automation & Cloud Engineering
- Design and implement automation frameworks to reduce operational overhead.
- Develop infrastructure automation using:
- Terraform
- CloudFormation
- AWS CDK
- Build operational automation using:
- Python
- Bash
- PowerShell
- AWS Lambda
- AWS Systems Manager
- Implement self-healing workflows and automated remediation mechanisms.
- Drive Infrastructure as Code adoption across managed environments.
Monitoring, Observability & Governance
- Architect and optimize monitoring and observability platforms using:
- Amazon CloudWatch
- Grafana
- Prometheus
- OpenSearch
- Tevico
- Establish monitoring standards and alerting strategies.
- Reduce alert fatigue through intelligent observability practices.
- Build operational dashboards and executive reporting frameworks.
Security & Compliance Operations
- Lead implementation of cloud security best practices.
- Manage and optimize:
- AWS Security Hub
- Amazon GuardDuty
- IAM Access Analyzer
- AWS Config
- CloudTrail
- Support compliance initiatives including:
- ISO 27001
- SOC 2
- HIPAA
- PCI DSS
- Conduct security posture reviews and remediation planning.
- Drive cloud governance initiatives across customer environments.
FinOps & Cloud Optimization
- Lead AWS cost optimization initiatives.
- Perform deep analysis of:
- Cost & Usage Reports
- Savings Plans
- Reserved Instances
- Resource Utilization
- Deliver actionable optimization recommendations.
- Participate in cloud financial governance reviews with customers.
- Drive continuous cost efficiency improvements.
Backup, Disaster Recovery & Business Continuity
- Design and validate backup and recovery strategies.
- Lead disaster recovery planning and testing exercises.
- Architect resilience solutions using AWS-native capabilities.
- Ensure compliance with recovery objectives (RTO/RPO).
- Conduct periodic recovery validation exercises.
Customer Advisory & Technical Leadership
- Act as a trusted technical advisor for managed services customers.
- Participate in customer operational reviews and governance meetings.
- Present recommendations related to:
- Reliability
- Security
- Cost Optimization
- Automation
- Performance Improvements
- Support solution reviews and operational architecture discussions.
- Collaborate with Presales and Cloud Engineering teams on customer engagements.
Team Leadership & Mentorship
- Mentor L1 and L2 CloudOps Engineers.
- Conduct technical reviews, operational coaching sessions & customer demos.
- Drive knowledge-sharing initiatives and capability development programs.
- Establish operational best practices and standards.
- Contribute to hiring and onboarding activities.
Documentation & Process Improvement
- Maintain and enhance:
- Runbooks
- SOPs
- Knowledge Bases
- Operational Playbooks
- Recovery Procedures
- Drive ITIL-aligned operational processes.
- Lead continuous service improvement initiatives.
- Contribute to Tevico feature feedback and operational enhancements.
Required Qualifications & Skills
- Bachelor's degree (B.E. / B. Tech.) in Computer Science, Information Technology, or related disciplines.
- 5–8 years of experience in Cloud Operations, Site Reliability Engineering (SRE), Infrastructure Engineering, DevOps, or Managed Services.
- Deep hands-on expertise in AWS cloud services.
- Strong experience managing production environments at scale.
- Advanced knowledge of:
- Linux Administration
- Windows Administration
- Networking
- DNS
- Load Balancing
- VPN
- Security Fundamentals
- Strong expertise in:
- Amazon EKS/ECS
- Amazon RDS
- Amazon S3
- IAM
- VPC
- Route 53
- AWS Backup
- CloudWatch
- Hands-on experience with:
- Terraform
- CloudFormation
- Infrastructure as Code
- Kubernetes core setup
- Amazon EKS/ECS
- Containers
- DevOps Toolchains
- Strong scripting skills using Python, Bash, or PowerShell.
- Experience leading incident response and RCA activities.
- Strong understanding of ITIL service management practices.
- Excellent troubleshooting, analytical, and communication skills.
Certification Requirements
Mandatory (At least two of the Following)
- AWS Certified Solutions Architect – Associate
- AWS Certified SysOps Administrator – Associate
- AWS Security Specialty
- ITIL Foundation Certification
Preferred
- AWS Certified DevOps Engineer – Professional
- AWS Solutions Architect – Professional
- AWS Advanced Networking Specialty
- Certified Kubernetes Administrator (CKA)
What We're Looking For
- Strong ownership mindset.
- Passion for automation and operational excellence.
- Customer-first attitude.
- Ability to remain calm under pressure.
- Strong mentoring and leadership capabilities.
- Continuous improvement mindset.
- High accountability and attention to detail.
What Success Looks Like
- Consistently maintains high availability and operational excellence across customer environments.
- Successfully leads major incident resolution and root cause analysis efforts.
- Delivers measurable improvements in automation, reliability, security, and cost optimization.
- Reduces operational toil through automation and process improvements.
- Acts as a trusted technical advisor to customers.
- Develops strong technical capability within the CloudOps team.
- Demonstrates readiness to progress into a Lead CloudOps Engineer role.
Why Join Comprinno
- Work with one of the leading AWS Managed Services and Consulting Partners in APJ.
- Manage large-scale enterprise cloud environments.
- Lead reliability, automation, FinOps, and cloud governance initiatives.
- Work closely with AWS experts, architects, and enterprise customers.
- Accelerate your career through advanced certifications and leadership opportunities.
- Contribute to Tevico, our cloud governance and observability platform.
- Join a culture that values innovation, operational excellence, ownership, and continuous learning.