Loading jobs…
Loading jobs…
Intermedia Cloud Communications
About Intermedia: Intermedia has established itself as a leading provider of cloud communications and collaboration tech that allows companies to connect better. We have a strong track record of growth, profitability, and creating an environment where everyone matters. Everyone.
While we are fast-paced and admittedly a bit intense, we promise that you won’t be bored. You will find Intermedia is a place where you can indulge your passion for creating and supporting great cloud technology. What’s more, we always look to promote from within and have many employees who have been with us 10, 15, and 20+ years!
Are you looking for a company where YOUR VOICE is heard? Where you can MAKE A DIFFERENCE? Do you THRIVE in a FAST-PACED work environment?
Do you wake every morning EXCITED to work with GREAT PEOPLE and create SUCCESS TOGETHER? Then Intermedia is the place for you. Culture at Intermedia is built on teamwork and transparency.
We hold each other accountable and always have each other’s back! Are you ready to make your mark?
About The Role
: While primarily remote , this role requires occasional visits to the office in Coimbra. We plan to open offices in Aveiro and Porto in the future. This approach gives team members the flexibility to work remotely while also coming together in the office for collaboration and teamwork.
Reliability Engineering Partner with Engineering teams to design resilient services, architectures, and deployment patterns. Define and promote SRE practices including SLIs, SLOs, error budgets, capacity planning, incident response, and post-incident learning. Identify systemic reliability risks and work with teams to address root causes.
Help reduce operational toil through automation, tooling, and better engineering practices. Architecture & Engineering Partnership Work actively with Engineering teams during design, development, and production-readiness reviews. Advise and challenge teams on service architecture, fault tolerance, scalability, observability, deployment safety, and operational readiness, helping them to make pragmatic trade-offs.
Support teams in diagnosing complex performance, latency, throughput, and resource-utilisation issues. Help establish engineering standards and reusable patterns for reliable, maintainable services. Performance & Observability Lead investigations into performance bottlenecks across applications, infrastructure, databases, queues, networks, and third-party dependencies.
Improve observability through metrics, logs, traces, dashboards, alerting, and service-level indicators. Help teams design meaningful alerts that identify user-impacting issues while reducing noise. Drive capacity planning and load-testing practices for critical systems.
Platform, Automation & Tooling Build and improve automation, deployment tooling, infrastructure-as-code, monitoring, and reliability platforms. Contribute to CI/CD improvements, release safety, rollback strategies, and progressive delivery practices. Develop tools that help Engineering teams self-serve reliability, diagnostics, and operational insights.
Improve cloud, container, and orchestration environments with a focus on security, reliability, and scalability. Incident Management & Operational Excellence Participate in incident response for high-priority production issues. Lead or contribute to blameless post-incident reviews.
Ensure actions from incidents result in improvements to architecture, tooling, monitoring, or process. Mentor engineers on production ownership and operational best practices. Experience in Site Reliability Engineering or senior backend/software engineering roles.
Software engineering background, with the ability to write clean, maintainable production code. Experience working with Engineering teams to influence architecture and improve production readiness. Understanding of distributed systems, scalability, resiliency patterns, failure modes, and performance engineering.