Site Reliability Engineer (SRE)

NEW

Site Reliability Engineer (SRE)

London, UK
Full-Time
$10,500/mo

About Us:

KAISPE is a technology company dedicated to delivering innovative digital solutions that help businesses improve efficiency, productivity, and growth. We combine modern technologies with industry expertise to build reliable, scalable, and high-quality solutions for our clients. Join our dynamic team and be part of a collaborative environment focused on innovation, reliability, and continuous improvement.

Job Description:

We are looking for a talented and experienced Site Reliability Engineer (SRE) to join our team in London. As an SRE at KAISPE, you will be responsible for ensuring the reliability, availability, scalability, and performance of our software applications and infrastructure. You will work closely with development, DevOps, QA, and product teams to build resilient systems, automate operational processes, and continuously improve the reliability of our technology platforms.

Key Responsibilities:

  • Design, implement, and maintain highly reliable and scalable infrastructure and systems.
  • Monitor system availability, performance, capacity, and overall reliability.
  • Develop automation tools and processes to reduce manual operational work.
  • Manage and optimize cloud infrastructure and deployment environments.
  • Implement monitoring, logging, alerting, and observability solutions.
  • Troubleshoot production incidents and identify root causes of system failures.
  • Participate in incident response, resolution, and post-incident reviews.
  • Collaborate with development teams to improve application performance and reliability.
  • Build and maintain CI/CD pipelines to support reliable and efficient software deployments.
  • Establish and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs).
  • Perform capacity planning and help ensure systems can scale with business requirements.
  • Implement backup, disaster recovery, and business continuity practices.
  • Document infrastructure, operational procedures, incident reports, and reliability standards.
  • Continuously evaluate and implement tools and technologies that improve system reliability and operational efficiency.

Requirements:

  • Bachelor’s degree in Computer Science, Software Engineering, Information Technology, or a related field.
  • Proven experience as a Site Reliability Engineer, DevOps Engineer, Cloud Engineer, or similar role.
  • Strong understanding of Linux/Unix systems and server administration.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud.
  • Proficiency in scripting or programming languages such as Python, Bash, or similar.
  • Strong understanding of networking, system administration, and distributed systems.
  • Experience with CI/CD pipelines and deployment automation.
  • Familiarity with monitoring, logging, and observability tools.
  • Experience with version control systems such as Git.
  • Strong troubleshooting, analytical, and problem-solving skills.
  • Understanding of containerization and orchestration technologies.
  • Ability to work effectively in a collaborative and fast-paced environment.
  • Good communication skills in English.

Preferred Qualifications:

  • Experience with Docker and Kubernetes.
  • Familiarity with infrastructure-as-code tools such as Terraform or Ansible.
  • Experience with monitoring and observability platforms such as Prometheus, Grafana, Datadog, or similar tools.
  • Knowledge of AWS services such as EC2, S3, RDS, CloudWatch, and IAM.
  • Experience with incident management and on-call operations.
  • Strong understanding of SRE principles, SLOs, SLIs, and error budgets.
  • Experience with security, networking, and cloud infrastructure best practices.
  • Familiarity with Agile/Scrum methodologies.
  • Experience working with high-availability and distributed systems.
  • Previous experience working in a fast-paced technology or software development environment.

What We Offer:

  • Competitive salary and benefits package.
  • Opportunity to work with modern cloud, infrastructure, and reliability technologies.
  • Collaborative and supportive working environment.
  • Flexible working hours and remote work options.
  • Professional development opportunities and career growth.
  • Health insurance and other employee benefits.
  • Opportunity to work with a talented and diverse technology team at KAISPE.
  • Exposure to large-scale systems, cloud infrastructure, automation, and DevOps practices.

How to Apply:

If you are passionate about building reliable systems, cloud infrastructure, automation, and improving technology operations, we would love to hear from you! Please send your resume and a cover letter to [your email address] with the subject line “Site Reliability Engineer (SRE) Application – [Your Name]”.

KAISPE is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.