Loading jobs…
Loading jobs…
Impact.Com — Victoria British Columbia
com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey. com empowers brands to drive trusted, performance-based growth through authentic relationships. Its award-winning products— Performance (affiliate), Creator (influencer), and Advocate (customer referral)—unify every type of partner into one integrated platform.
com helps brands show up where it matters most. com to power more than 225,000 partnerships that deliver measurable business results. com: As a Site Reliability Engineer, you'll be the champion of performance and stability for our core application ecosystem.
Working closely with our Java and C# engineering squads, you'll ensure that our high-frequency data ingestion pipelines and customer-facing applications meet strict performance and error rate benchmarks. Your mission is to bridge the gap between code and infrastructure, implementing OpenTelemetry (OTel) standards across the stack to provide deep visibility into how we interact with external social APIs and how our internal services communicate.
What You'Ll Do
: OTel Orchestration: Become the architect of our observability pipeline. Implement and maintain OpenTelemetry instrumentation across Java and C# services to ensure high-fidelity traces, metrics, and logs. API Reliability: Build integration tests with third-party social APIs and setup the appropriate monitoring and alerting systems to ensure high availability and reliability.
NET environments. Infrastructure as Code: Drive root-cause analysis (RCA) for complex distributed system failures and contribute to remediations through code optimizations or infrastructure adjustments. Distributed Tracing: Leverage tracing data to identify bottlenecks in cross-service communication and optimize the path of data from social APIs to our internal stores.
Full-Stack Troubleshooting: Debug issues across the entire stack, from containerized application code (Java/C#) down to network calls and cloud resource utilization. Capacity Planning: Analyze application usage patterns to inform scaling decisions, ensuring we handle social data bursts without compromising stability or overspending on cloud costs. What You Bring: Software Pedigree: Strong proficiency in Java or C#.
You are comfortable reading, debugging, and instrumenting application code. Observability Expert: Hands-on experience with OpenTelemetry, including auto-instrumentation, manual spans, and collector configuration. Modern Tooling: Deep experience with the Grafana ecosystem (Prometheus, Tempo, Loki) or similar distributed tracing platforms (Jaeger, Honeycomb, Datadog).
API Savvy: Experience working with high-volume REST/Graph APIs and an understanding of OAuth flows, rate-limiting, and webhooks. Systems Mindset: Solid understanding of how Java/C# applications interact with the underlying infrastructure. Pragmatism: Ability to prioritize tasks in a high-velocity environment and a focus on building "self-healing" systems rather than manual fixes.
S. in Computer Science, or equivalent practical experience in a high-scale production environment.
Nice To Have
: Affiliate Partnerships Industry Fundamentals Certification by PXA Salary Range: $110,000 - $130,000 per year, plus additional 5% variable annual bonus contingent on Company performance and eligible to receive Restricted Stock U