Site Reliability Engineering is an engineering discipline that combines software and systems engineering to
build and run systems. SRE ensures that Aswat s services both our internally critical and our
externally-visible systems have reliability and uptime appropriate to customers needs and a fast rate of
improvement while keeping an ever-watchful eye on capacity and performance.
SRE is also a mindset and a set of engineering approaches to running better production systems we build
our own creative engineering solutions to operations problems. Much of our software development
focuses on optimizing existing systems, building infrastructure and eliminating work through automation.
As SREs are responsible for the big picture of how our systems relate to each other, we use a breadth of
tools and approaches to solve a broad spectrum of problems. Practices such as limiting time spent on
operational work, blameless postmortems and proactive identification of potential outages factor into
iterative improvement that is key to both product quality and interesting and dynamic day-to-day work.
SRE s culture of diversity, intellectual curiosity, problem solving and openness is key to its success. Our
organization brings together people with a wide variety of backgrounds, experiences and perspectives. We
encourage them to collaborate, think big and take risks in a blame-free environment. We promote
self-direction to work on meaningful projects, while we also strive to create an environment that provides
the support and mentorship needed to learn and grow.
Role And Responsibilities
Engage in and improve the whole lifecycle of services from inception and design, through
deployment, operation and refinement.
Support services before they go live through activities such as system design consulting,
developing software platforms and frameworks, capacity planning and launch reviews.
Maintain services once they are live by measuring and monitoring availability, latency and
overall system health.
Scale systems sustainably through mechanisms like automation, and evolve systems by pushing
for changes that improve reliability and velocity.
Practice sustainable incident response and blameless postmortems.
BS degree in Computer Science or equivalent practical experience.
Experience with algorithms, data structures, complexity analysis and software design.
Required Skills (Required To Apply)
Systematic problem-solving approach, coupled with strong communication skills and a sense of
ownership and drive.
Ability to debug and optimize code and automate routine tasks.
Experience building APIs and data delivery mechanisms for applications such as web
dashboards, alerting & health systems, mobile applications, and third party integration
Ability to design and manage CI / CD pipelines
Strong Knowledge of:
Git & Git Flows
Additional Skills (Big Advantage)
FreeSwitch or Asterisk