Aditya

Aditya

I am a determined and diligent professional, driven by a deep ambition to achieve mastery through focused, logical analysis. My structured approach and intense concentration enable me to meticulously evaluate and benchmark complex AI systems, ensuring precision and quality in every detail.

Open to

Bangalore · Hyderabad · Gurgaon

Work style

remote

Contact details — on requestProof of Work — on request

Experience

M

AI Evaluation, Data Annotation & Benchmarking Subcontracts

Multiple Frontier AI Companies · Present

Not yet confirmed
  • Worked across AI data annotation, dataset creation, RLHF/SFT data generation, expert model evaluation, calibration, rubric creation, task validation, and quality assurance for frontier AI systems.
  • Created and evaluated repository-level software-engineering tasks across Python, Go, and Rust, including backend systems, APIs, games, debugging, and infrastructure-oriented problems across repositories such as etcd, Kubernetes, Vault, and Traefik.
  • Performed expert RLHF / preference evaluation of model-generated patches for highly technical repository-level tasks, assessing correctness, implementation quality, regressions, edge cases, and issue resolution using detailed rubrics and test cases.
  • Worked on highly challenging SWE-bench-style repository problems for frontier models, selecting difficult issues from areas of expertise and defining task requirements, expected behavior, rubrics, and validation tests for broken or incomplete model-generated solutions.
  • Created 60+ terminal-based benchmark tasks across mathematics, formal methods, scientific computing, software engineering, code debugging, reverse engineering, low-level programming, business knowledge, and quantitative reasoning.
  • Contributed to CI/CD and evaluation-harness engineering, supporting reproducible task execution, environment setup, trajectory capture, automated verification, test execution, and model-output evaluation.
  • Created and reviewed a large volume of computer-use tasks across Linux, Windows, and macOS, covering GUI workflows, application interaction, file manipulation, browser workflows, and multi-step real-world computer operations.
  • Worked across RL and SFT data pipelines, including task creation, data generation, preference evaluation, calibration, reviewer workflows, and quality validation for computer-use agents.
  • Worked with the tooling team on environment and evaluation infrastructure across Linux, Windows, and macOS, contributing to task execution, evaluator workflows, benchmark reliability, and cross-platform automation.
  • Took on tooling-engineering and forward-deployed engineering responsibilities, contributing to workflow optimization, evaluation efficiency, automation, and benchmark cost-reduction initiatives.
  • Automated generation and validation of task JSON specifications, producing 90-100% instruction/evaluator alignment across generated task specifications and reducing manual effort in benchmark preparation.
  • Created benchmark tasks designed around defined single-agent and multi-agent performance thresholds to evaluate frontier systems on complex task execution and multi-agent coordination.
  • Took on lead and reviewer responsibilities, reviewing and calibrating benchmark tasks and validating task specifications, evaluation criteria, quality, and scoring consistency.
  • Created and evaluated 10 unique long-horizon, greenfield repository tasks, targeting difficult software-engineering scenarios and measurable performance separation between evaluated coding agents.
  • Designed high-difficulty tasks targeting a measurable performance separation of approximately 20% between evaluated systems while maintaining challenging, realistic repository-generation requirements.
P

Senior Software Developer

Pyvision tech · Jan 2026 – Present

Not yet confirmed
  • • Built a trading and backtesting platform using Python, FastAPI, Supabase, and NSE APIs
  • • Implemented scalable backtesting pipelines using backtesting.py
  • • Integrated vector databases and developed AI orchestration layers using MCP frameworks
  • • Worked on LLM evaluation pipelines including SWE bench style workflows, contributing to dataset creation and validation systems
  • • Promoted to Team Lead, conducted reviews and handled QA processes to ensure quality and consistency across tasks
P

Senior Software Engineer (AI)

Pyvision Technologies · Mar 2025 – Present

Not yet confirmed
  • Led engineering initiatives through code reviews, architecture decisions, technical planning, delivery ownership, and engineering-quality validation across software projects.
  • Built a finance-oriented trading and backtesting platform using Python, FastAPI, Supabase, NSE APIs, vector databases, and data pipelines for financial-data ingestion, transformation, analysis, and backtesting workflows.
  • Designed scalable backend services and APIs using Python and FastAPI, with containerized development and deployment workflows using Docker, Kubernetes, Linux, and CI/CD.
  • Developed production-oriented Shopify systems using Next.js, TypeScript, React, Python, and FastAPI, integrating frontend applications with backend APIs and data services.
  • Developed analytics dashboards using React, Next.js, and Plotly for operational visibility, auditing, business metrics, and system monitoring.
  • Integrated vector databases, RAG pipelines, LangChain, LlamaIndex, MCP frameworks, OpenAI APIs, and Claude APIs to build context-aware retrieval and AI-powered tool-orchestration workflows.
  • Designed agentic workflows connecting LLMs, tools, APIs, data sources, and backend services for multi-step automation and AI-powered product features.
  • Implemented structured prompt engineering, few-shot prompting, and multi-step reasoning workflows for LLM-powered product features and Al systems.
  • Built Docker-based isolated environments and testing workflows for SFT/RLHF and LLM evaluation scenarios, supporting reproducible execution and model-quality analysis.
  • Managed containerized deployments using Docker and Kubernetes, including CI/CD automation, Linux-based environments, deployment workflows, testing, and monitoring.
P

Software Developer

Pyvision Tech · Mar 2025 – Jan 2026

Not yet confirmed
  • • Built production grade frontend systems for Shopify platforms using Next.js and TypeScript
  • • Developed backend systems and APIs using Python and FastAPI
  • • Worked with Docker based environments for development and testing workflows
  • • Contributed to agent environments such as OSWorld and terminal bench based setups, focusing on real world task execution across GUI and CLI workflows
  • • Worked with repository level tasks and multi step execution workflows.
O

AI Evaluation / RLHF Contractor (Freelance)

Outlier · Aug 2024 – Apr 2025

Not yet confirmed
  • Created 40+ production-quality frontend implementations using HTML/CSS/JavaScript, TypeScript, React, and Next.js, extending model-generated interfaces into realistic end-user websites.
  • Implemented interfaces across minimalist, 2D, 3D, flat, glassmorphic, and skeuomorphic design patterns, including wireframing for selected tasks and detailed visual and interaction requirements.
  • Developed high-quality golden/reference solutions inspired by products such as Tesla, Apple, Zapier, and Discord, enabling direct comparison and quality evaluation of model-generated frontend outputs.
  • Reviewed trainer-created frontend solutions against task requirements and reference implementations, identifying functional, visual, implementation, and instruction-following deficiencies.
  • Performed extensive multi-turn RLHF, preference annotation, response ranking, and model-output evaluation across technical model outputs, with the majority of work focused on software engineering, coding, reasoning, computer-use, and agentic workflows.
  • Built interactive data-visualization dashboards using Plotly and performed multi-turn preference evaluation to rank, compare, and improve model outputs across iterative interactions.
  • Reviewed and calibrated model outputs across data-visualization and other technical tasks, providing structured preference feedback, quality assessments, consistency checks, and reviewer-level validation.
  • Created and evaluated training data, annotations, task specifications, reference solutions, rubrics, and validation criteria for improving model performance on complex technical tasks.
  • Performed RLHF and model-output evaluation across diverse task categories, including approximately 90% technical and 10% non-technical evaluation work.
  • Worked on GitHub repository tasks involving bug fixes, backend development, production-oriented changes, and code refactoring, while defining detailed rubrics and expected outcomes for model evaluation.
  • Worked with containerized GitHub repositories to reproduce tasks against controlled code versions, validate model-generated patches, and establish deterministic execution conditions for coding-agent evaluation.
  • Evaluated repository-level solutions by comparing model-generated patches against expected behavior, test outcomes, implementation requirements, and task-specific evaluation criteria.
  • Created structured reference solutions and validation artifacts to provide models with clear target implementations, expected outcomes, and evaluation signals for technical software-engineering tasks.
R

Software Developer Intern

Revino · Jan 2025 – Mar 2025

Not yet confirmed
  • * Built core frontend and backend features for a full stack product using React, Node.js, Express, and MongoDB
  • * Developed REST APIs and worked on integration between frontend and backend modules
  • * Contributed in an agile team environment with Git based workflows, testing and feature delivery
  • * Worked on bug fixes, feature improvements and overall product development during the internship

Skills 0 proven through work

Also works with

UI/UX Design Tools UsageRaft Algorithm ImplementationETCD UsageBacktesting.py UsageVector BT UsageEffective ContributionSoftware Architecture PrinciplesOn-time DeliveryPersistenceLoss Function OptimizationMathematicsAPI Performance OptimizationModel Fine-TuningJSON Data HandlingWorkflow Automation (Spreadsheet-Based)Workflow CoordinationArchitectural DesignTeam LeadershipAI Model ConfigurationTest Case EvaluationModel Training ExecutionReward Function DesignReinforcement LearningResearch WritingProduct DevelopmentSoftware Development PracticesTraining Data CreationSEO Strategy & RoadmappingUser-Centered DesignUser Interface DesignMinimalist Design PrinciplesUI Animation DevelopmentForm Validation ImplementationCustom Component DevelopmentWireframingFront-end DevelopmentBuilding Web ApplicationsNext.js DevelopmentTypeScript DevelopmentReact DevelopmentJavaScript DevelopmentCSSHTMLAWS DeploymentGoogle Cloud PlatformData Flow DesignPrompt EngineeringTechnical TroubleshootingDistributed Systems DesignKubernetes OrchestrationForward-Deployed EngineeringProblem SolvingModel AdaptationCron Job SchedulingDatabase Query OptimizationPerformance TestingTechnical IndicatorsMutual Fund AnalysisFinancial Market AnalysisAPI IntegrationInformation ResearchModel Performance ImprovementBenchmarking ExecutionTool DevelopmentReinforcement Learning Environment CreationScalable System DesignCode RefactoringBug FixingMathematical ModelingDocument ParsingData TransformationText ClassificationGamification ImplementationFeature IntegrationVector Database UsageTime Series Trend AnalysisReal-Time Data IntegrationPattern RecognitionFinancial AnalysisAlgorithmic TradingTrading Platform Development

Proof of Work

Proof of Work

Aditya shares this with people who ask. You'll hear back either way.

Education

Bachelor of Technology, Computer Science and Engineering

Jawaharlal Nehru Technological University, Hyderabad · 2019 — 2023

Contact details

Contact details

Aditya shares this with people who ask. You'll hear back either way.

What drives their work

The Inventor — Makes what did not exist

The Inventor

They envision new challenges and build the systems to test them.

Aditya's superpower

You excel at designing and implementing novel, complex environments and benchmarks to push the limits of AI models.

How Aditya works

You are best suited for highly technical, research-oriented teams focused on advancing AI capabilities and creating novel solutions.

Ask about this candidate

AI responses may contain errors. Proof status and source information are provided by Proof of Skill.

A Proof CV — one profile, kept current, with the proof attached. Back to the portal