Site Reliability Engineer

Job Description:
Title: Site Reliability Engineer
Employment type: Long term contract
Location: London, UK
Mode: 5 days onsite
Domain: Banking
As a Site Reliability Engineer (SRE), you'll help build a meaningful engineering discipline, combining software and systems to develop creative engineering solutions to operations problems. Much of our support and software development focuses on optimizing existing systems, building infrastructure, and reducing work through automation. You’ll join a team of curious problem solvers with a diverse set of perspectives who are thinking big and taking risks. In this environment, you’ll take the lead on relevant projects, supported by an organization that provides the support and mentorship you need to learn and grow. As an SRE, you’ll be focused on running better production applications and systems.
Job Responsibilities
Design, code, test, and deliver software to automate manual operational work
Troubleshoot priority incidents, facilitate blameless post-mortems and ensure permanent closure of incidents
Engage with development team throughout the life cycle to help develop software for reliability and scale, ensuring minimal refactoring or changes
Identify application patterns and analytics in support of better service level objectives
Design self-healing and resiliency patterns
Design automated software and product upgrades, change management, and release management solutions
Coach or manage teams as applicable
Participate in the 24x7 support coverage as needed
Required qualifications, capabilities, and skills
Bachelor’s degree or equivalent experience in software engineering discipline
Exposure to tools related to CI/CD, Application Resiliency, and Security
Exposure to microservice architecture, AWS and Containers
Strong Python needed.
Emerging knowledge of software applications and technical processes within a technical discipline, specifically building out cloud native solutions (AWS).
Working knowledge of infrastructure components (e.g. routers, load balancers, cloud products, container systems, compute, storage, and networks)
Excellent debugging and trouble shooting skills.