Wipro Limited (NYSE: WIT, BSE: 507685, NSE: WIPRO) is a leading technology services and consulting company focused on building innovative solutions that address clients’ most complex digital transformation needs. Leveraging our holistic portfolio of capabilities in consulting, design, engineering, and operations, we help clients realize their boldest ambitions and build future-ready, sustainable businesses. With over 230,000 employees and business partners across 65 countries, we deliver on the promise of helping our customers, colleagues, and communities thrive in an ever-changing world. For additional information, visit us at www.wipro.com.
Job Description
Role Overview As an AI Data Foundry, we partner with top-tier labs to build the critical SFT and RLHF datasets that train the next generation of LLMs, SLMs, and Physical AI. We are seeking a Lead Researcher who focuses on the rigorous science of measurement and alignment. In this role, you will architect preference data pipelines, design proactive benchmarks from the ground up, and publish novel methodologies on arXiv and at top-tier conferences (e.g., NeurIPS, ICLR).
͏
Key Responsibilities
● Design Alignment Methodologies: Architect frameworks for human-in-the-loop and synthetic data generation. Design pairwise preference rubrics (utilizing Bradley-Terry models) and reward mechanisms for DPO, KTO, and PPO pipelines.
● Build Proactive Benchmarks: Lead the creation of novel, domain-specific benchmarks, spanning agentic reasoning to physical AI, that evaluate complex capabilities rather than relying on saturated, static datasets.
● Decontaminate & Validate: Build robust pipelines to detect data contamination (via n-gram overlap, embedding similarity) to ensure deliverables are mathematically sound and mitigate reward hacking.
● Research Publication: Conduct independent research on evaluation frameworks and alignment science, co-authoring papers to establish our technical authority in the field.
● Technical Leadership: Work closely with domain-specific researchers to synthesize ground-truth data into cohesive, scoreable, and statistically sound evaluation frameworks.
͏
Candidate Profile
● Academic Background: Advanced degree (PhD or high-impact MSc) in Computer Science, Artificial Intelligence, Mathematics, or a highly quantitative field.
● Technical Expertise: Deep familiarity with modern alignment techniques (RLHF, DPO, RLAIF) and standard evaluation harnesses (e.g., EleutherAI LM Eval, HELM, OpenAI Evals).
● Core Stack: Advanced proficiency in PyTorch, HuggingFace, vLLM, and complex data manipulation.
● Model Experience: Hands-on experience fine-tuning or evaluating open-weight models (Llama 3, Mistral, Qwen) with a deep understanding of their architectural bottlenecks.
͏
͏
Expected annual pay for this role ranges from $240,000.00 to $375,000.00. Based on the position, the role is also eligible for Wipro’s standard benefits including a full range of medical and dental benefits options, disability insurance, paid time off (inclusive of sick leave), other paid and unpaid leave options
Reinvent your world. We are building a modern Wipro. We are an end-to-end digital transformation partner with the boldest ambitions. To realize them, we need people inspired by reinvention. Of yourself, your career, and your skills. We want to see the constant evolution of our business and our industry. It has always been in our DNA - as the world around us changes, so do we. Join a business powered by purpose and a place that empowers you to design your own reinvention.
Applications from people with disabilities are explicitly welcome.