Job Description
Role Overview
Help build and scale the tooling we use to measure how well AI-powered software development tools actually perform. You'll develop evaluation harnesses, automate benchmark runs, and help make sure the results we produce are reproducible and hold up to scrutiny. This is an engineering role, but a lot of the work is about getting the measurement right, not just automating it.
͏
Key Responsibilities
• Build and integrate evaluation harnesses and automation for software development use cases, including turning real engineering artifacts like merged pull requests into repeatable benchmark tasks.
• Build versioned, repeatable processes to evaluate AI tools, models, and harnesses, with reproducible run environments (pinned dependencies, containerized runs, isolated worktrees) so results stay comparable over time.
• Validate and calibrate evaluation approaches against human judgment, so scores are consistent and correct rather than just repeatable.
• Support execution-based benchmarking across quality, productivity, and e iciency measures, including cost and latency.
• Analyze results across repeated runs, looking at variance, failure patterns, and cost per outcome, and find ways to make the workflows more reliable and more automated.
• Work with engineering and data teams to improve the tooling, and document how the evaluations work and what they found for both technical and leadership audiences.
͏
Required Skills & Experience
• Strong software engineering background, with real experience building automation, developer tooling, or test and validation systems.
• Coding-agent evaluation: Evaluating AI coding agents that modify code repositories, including validating generated code changes against expected outcomes.
• Evaluation harnesses: Building automated, reproducible evaluation workflows, including test execution, environment setup, and result validation.
• Benchmarking and reliability: Establishing baselines, measuring run-to-run variance, analyzing failures, and ensuring consistent evaluation results.
• Git and CI/CD integration: Working with repository history, branches, pull requests, and automated testing within CI/CD pipelines.
• LLM-as-a-Judge: Using LLMs to evaluate coding-agent outputs, including calibration against human assessments.
• Proficient in at least one general-purpose language such as Python, Java, or JavaScript — the specific language background is flexible.
• Solid working knowledge of Git, including how branches, history, and working trees behave, and of containerization with Docker.
• Experience with APIs, development environments, CI/CD pipelines, and typical engineering workflows.
• Understanding of how AI, LLM, or agent evaluation works and where it goes wrong, such as why a judge can be consistent but still wrong, why a single run can mislead, and how benchmark contamination happens.
• Able to troubleshoot technical problems, think clearly about whether a measurement is valid, and analyze results carefully.
• Hands-on experience using AI coding tools and agentic harnesses such as Claude Code, Devin, or Cursor, and command of the best practices for working with them effectively.
Mandatory Skills: Cloud Product & Platform Testing .
Experience: 5-8 Years .
The expected compensation for this role ranges from $60,000 to $148,500 .
Final compensation will depend on various factors, including your geographical location, minimum wage obligations, skills, and relevant experience. Based on the position, the role is also eligible for Wipro's standard benefits including a full range of medical and dental benefits options, disability insurance, paid time off (inclusive of sick leave), other paid and unpaid leave options.
Applicants are advised that employment in some roles may be conditioned on successful completion of a post-offer drug screening, subject to applicable state law.
Wipro provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws. Applications from veterans and people with disabilities are explicitly welcome.
Reinvent your world. We are building a modern Wipro. We are an end-to-end digital transformation partner with the boldest ambitions. To realize them, we need people inspired by reinvention. Of yourself, your career, and your skills. We want to see the constant evolution of our business and our industry. It has always been in our DNA - as the world around us changes, so do we. Join a business powered by purpose and a place that empowers you to design your own reinvention.