What you'll do: quality and signal at enterprise scale
- As a Software Engineer on our Quality & Observability team, you will own the testing and telemetry that let us ship a major platform release with confidence. Our systems are distributed and event-driven, built on containerized services with eventual consistency at their core. You will design the functional test coverage and the observability instrumentation that prove those systems behave correctly and stay measurable in production. This is hands-on engineering across a heterogeneous, multi-service estate, guided by architecture and standards rather than by any single codebase.
In this role, you will:
- Design and build functional and contract test assets for cloud-native services, asserting on real system state and side effects, not just on HTTP response codes
- Instrument services end to end with OpenTelemetry, producing traces and metrics that export cleanly and remain usable for real operational monitoring
- Derive test cases and telemetry points directly from architecture documentation, ensuring every documented invariant has a test and every named span or metric is actually observed
- Build test patterns for distributed, asynchronous systems that respect eventual consistency, readiness gating, and lifecycle signals rather than assuming synchronous request and response
- Design representative test data and fixtures for a complex domain, including patterns for validating probabilistic and AI-generated outputs against thresholds and golden datasets
- Work independently across many repositories, following documented conventions with minimal hand-holding and contributing improvements back to shared tooling and CI workflows
- Collaborate with product owners, architects, and cross-functional teams to deliver release-critical quality and telemetry work on schedule
- What makes you a qualified candidate
Skills in action
- We are seeking engineers with a strong foundation in DevOps practices, containerization, and the realities of testing distributed systems. You bring:
- Strong proficiency in Python and a modern test framework such as pytest, comfortable writing tests to a repository's established conventions
- Hands-on experience with Docker and Docker Compose, able to read and extend container definitions to compose real service dependencies locally rather than mocking them
- Solid literacy in distributed and event-driven systems, including message-driven architectures, eventual consistency, and asynchronous processing patterns
- Experience with API and contract testing against HTTP services, including modern Python frameworks such as FastAPI
- Familiarity with validating side effects across data stores such as relational databases, search indexes, and graph or NoSQL systems after a change or mutation
- Experience testing message-driven flows: producing test events, asserting consumer-side effects, and reasoning about partitioning and ordering semantics
- Comfort designing assertions for probabilistic or model-generated output using threshold-based and golden-dataset patterns rather than exact matching
- Disciplined Git and GitHub workflow habits, including branch strategy and pull-request review gates
- The ability to work from an architecture document, deriving test cases and instrumentation points from documented behavior and invariants
Job Description
How to read:
The job description below should be read in conjunction with the purpose of the job family.
Job Title: Senior Executive âÃÂàJob Family
<Insert Job Family definition here>
Job Purpose: The purpose of the role is timely completion of assigned tasks with quality and work on areas of improvement.
The role involves implementing process efficiencies and improving outcomes. The position requires collaboration with teams and stakeholders to ensure accurate, timely deliverables while upholding operational and quality standards
͏
Job Duties and Responsibilities:
- Ensure timely completion of tasks to meet project and functional objectives.
- Implement tools or methods as shared by the organization to streamline operations and minimize delays within their area of work.
- Ensure compliance with regulatory standards as applicable.
- Adhere to quality standards and contribute to improve them
- Prepare documentation, reports, or presentations, as per quality standards.
- Contribute to new ideas and process improvements.
- Proactively resolve operational issues to ensure continuity and minimal disruptions.
- Collaborate with leads and peers to ensure tasks are completed on time
- Coordinate effectively within team to enhance communication and outcomes.
- Participate in training programs to enhance expertise and share knowledge to support team development.
͏
͏
͏
Observability and instrumentation
- Telemetry expertise
- This role has a strong observability focus. Ideal candidates bring:
- Hands-on experience with the OpenTelemetry SDK, instrumenting services for traces and metrics with OTLP export
- Experience instrumenting HTTP services and agent or worker components, using auto-instrumentation libraries where available
- Understanding of trace-context propagation across service and asynchronous message boundaries, not just HTTP header propagation
- Familiarity with observability backends such as Grafana, Loki, Tempo, and Mimir, useful for verifying that instrumentation produces genuinely usable signal
- A config-driven approach to instrumentation that respects environment-based configuration and avoids hardcoded endpoints
- Bonus points
Nice to have
- Experience with GitHub Actions or comparable CI pipelines, including shared or reusable workflows
- Exposure to OpenTelemetry instrumentation in Rust, using the tracing ecosystem
- Prior work with knowledge graphs, data catalogs, or metadata-management platforms
- Bachelor's or Master's degree in Computer Science, Engineering, Data Science, or a related field, or equivalent practical experience