DevOps Engineer
Software Engineering
Hyderabad, Telangana, India
Job Description
DevOps Engineer (Mid–Senior) – Job Description
About Apty
Apty is a Digital Adoption Platform (DAP) that operates directly within web applications to improve user productivity, enforce process compliance, and eliminate friction in complex enterprise workflows.
Technically, Apty sits in a unique layer: on top of third-party applications, requiring deep runtime integration, browser extensions, and a high-scale multi-tenant backend processing millions of user interactions.
This is systems-heavy engineering, not standard infrastructure maintenance.
About the Role
We are looking for a highly hands-on Mid–Senior DevOps Engineer who can take strong ownership of our DevOps ecosystem and grow into owning it end-to-end.
With a growing base of 2M+ active users, our platform requires high availability, real-time processing, and extreme scalability.
This role is not about maintaining pipelines—it’s about:
- Architecting cloud-native distributed systems
- Scaling high-throughput analytics pipelines
- Building enterprise-grade SaaS and on-prem deployments
- Driving DevOps excellence across the company
If you enjoy owning infrastructure like a product, this role will suit you well.
What You’ll Work On in Your First 90 Days
- Take ownership of our AWS infrastructure and CI/CD ecosystem
- Optimize and scale event ingestion pipelines
- Strengthen observability, alerting, and system reliability
- Improve deployment workflows across microservices, extensions, and frontend applications
- Contribute to architecture decisions for scaling to millions of users
- Start driving on-prem deployment standardization
Requirements
Key Responsibilities
Cloud Infrastructure & Architecture (AWS)
- Architect and manage scalable infrastructure using AWS services
- Design and maintain a multi-tenant SaaS architecture
- Scale systems handling millions of real-time events
- Optimize performance, cost, and reliability
DevOps & Automation
- Own and enhance CI/CD pipelines using GitHub Actions
- Build automated deployments for:
- Microservices
- Backend systems
- UI applications
- Browser extensions
- Microservices
- Implement Infrastructure-as-Code using Terraform
- Build and maintain Dockerized environments
- Improve release velocity while maintaining stability
Monitoring, Security & Reliability
- Implement end-to-end observability across logs, metrics, and traces
- Set up alerting and incident-response systems
- Ensure:
- High availability
- Auto-scaling
- Disaster recovery
- Backup strategies
- High availability
- Drive security best practices, including:
- Network isolation
- IAM policies
- CVE remediation
- Container security
- Network isolation
On-Prem Deployments
- Design and deliver on-premise versions of our SaaS platform
- Replicate cloud architecture in constrained enterprise environments
- Build automation for:
- Installation
- Upgrades
- Maintenance
- Installation
- Collaborate with customers to customize deployments
Collaboration & Leadership
- Work closely with Engineering, QA, and Product teams
- Share knowledge and support fellow engineers on DevOps practices
- Help drive best practices across teams
- Contribute to architecture and design decisions
- Act as a technical owner for DevOps initiatives
Required Skills & Qualifications
Core Expertise
- 3+ years in DevOps, with hands-on ownership of production systems
- Strong expertise in AWS, which is mandatory
- Hands-on experience with:
- Kubernetes (EKS)
- VPC, EC2, RDS, and ElastiCache
- S3, CloudFront, and Route 53
- API Gateway, Lambda, and Kinesis
- ClickHouse or similar columnar databases
- Docker and container orchestration
- Helm charts
- Kubernetes (EKS)
CI/CD & Infrastructure
- Strong experience with GitHub Actions
- Expertise in Terraform, which is mandatory
Programming and Scripting
Strong scripting skills in:
- Python
- Shell
- Node.js
Development experience is a strong advantage.
Scalability & Distributed Systems
- Experience handling high-scale distributed systems
- Strong understanding of:
- Load balancing
- Caching strategies
- Event-driven architectures
- High-throughput pipelines
- Load balancing
On-Prem & Enterprise Systems
- Experience building SaaS-to-on-prem deployments
- Ability to automate complex enterprise setups
Other Requirements
- Strong debugging and problem-solving skills
- Ownership mindset—drives problems end-to-end
- Excellent communication and collaboration skills
Nice-to-Have Skills
- Experience with Digital Adoption Platforms or browser-based products
- Experience with monitoring tools such as:
- Grafana
- Prometheus
- Datadog
- Grafana
- Exposure to:
- SOC 2
- ISO compliance environments
- SOC 2
What Makes You Stand Out
- You treat infrastructure as a product, not a support function
- You think in systems, trade-offs, and scale
- You can design from scratch, not just maintain
- You are comfortable handling high ambiguity and ownership
- You optimize for long-term reliability, not short-term fixes
Interview Process
- Initial screening: 15 minutes
- Technical deep dive: 60 minutes
- DevOps/system design round: 60 minutes
- Leadership discussion: 30–45 minutes
End-to-end timeline: Approximately two weeks
Interested?
Contact: hr@apty.ai