Surbhi Choudhary

Surbhi Choudhary

Open to

Bangalore · Mumbai · Singapore · Other locations

Work style

onsite, hybrid, remote

Contact details — on requestProof of Work — on request

Experience

Capgemini

Senior AI ML Engineer

Capgemini · May 2025 – Present

Not yet confirmed
Capgemini Technology Services India Limited

Senior Consultant- AI Product Lead

Capgemini Technology Services India Limited · May 2025 – Present

Not yet confirmed
  • Architected and owned end-to-end product delivery from business requirement discovery, architecture definition, commercial estimation, technical implementation, production deployment, stakeholder demonstrations, and continuous platform enhancement.
  • MCP-based unified Agentic AI Orchestration Platform integrates diverse multimodal applications including text-to-audio, video-to-video, and doc-to-doc systems-into a single interface with multi-step workflows using LangGraph and LangChain to facilitate seamless data handoffs between independent AI modules and specialized media processing tools enabling a modular and scalable architecture for cross-functional automation, serving stakeholders across 80+ countries and 60+ languages.
  • Dashboard integrated measuring operational, functional, non-functional evaluation measures, saved 2 billion dollars annually.
  • Led and architected end-to-end delivery of enterprise multilingual LLM- based AI platform on Azure including stakeholder management, development, and maintenance supporting both individual systems text-to-audio and video-to-video transformation with accent preservation, native tone preservation, white segments, spectrogram analysis using Huggingface Coqui TTS, Whisper SST, MarianMT models for on-premises and Azure Open AI and Azure Cognitive Services for cloud version.
  • Developed a Streamlit-based application that analyzes text and associated audio files using LangDetect for language detection and Whisper for transcription, identifying foreign words, and detecting mismatches between text and audio content, later integrated in main application used to detect any language mismatch and foreign words detection as fallback mechanism.
  • Built a full-stack POC using OpenAI model with Python, AWS S3, and MongoDB to generate phonemes and audio from user input words/languages, enabling users to select preferred phonemes to create a pronunciation library for main system.
  • Architected a hybrid Document Translation Engine for doc-to-doc translation leveraging T5 model for local version, and open AI model for privacy-focused processing and for high-complexity linguistic tasks preserves original document structures, including layouts, formatting and tables, during the translation of sensitive enterprise files.
  • Engineered a local-first deployment to ensure 100% data security, eliminating the need for cloud storage while maintaining high-speed document automation.
  • Separate modules designed for files with unstructured formats of tables and OCR for data extraction from scanned images.
  • Led a 5-member team to design and deliver RAG-based AIOps solutions for ServiceNow, enhancing analytics through Azure data pipelines, embeddings, and AI-powered search integrated with Microsoft Copilot.
  • Partnered with product owners and enterprise stakeholders to prioritize capabilities, define delivery milestones, and balance technical implementation with operational business objectives.
  • Collaborated with AI solution architects to evaluate and standardize technology selections, deployment architectures, and platform components based on scalability, security, maintainability, enterprise constraints, and long-term sustainability.
  • End-to-end system ingesting ServiceNow incident and change request data into Azure Blob Storage, creating embeddings and Azure AI Search indexes to enable keyword and semantic search, frequency analysis, and advanced date-based filtering.
  • Automated incident resolution recommendations by leveraging historical root cause analysis (RCA) documents, enabling the application to intelligently correlate new incidents with past patterns and suggest probable causes and remediation actions.
  • Drove 70%+ performance improvement in the AIOps ServiceNow Incident Intelligence platform through Azure AI Search hybrid indexing and LLM-powered RCA synthesis - reducing mean time to resolution for IT incident management, enabled support teams to reduce manual troubleshooting effort and accelerate incident resolution.
  • Collaborated with business stakeholders, product owners, and enterprise leadership to prioritize AI initiatives based on business impact, implementation feasibility, security, and long-term organizational value and to translate regulatory, compliance, and operational objectives into explainable AI-driven risk assessment workflows and decision-support capabilities.
  • Designed and led development of an AI-driven risk assessment platform using adaptive, LLM-powered questionnaires to evaluate application's risk factors across domains such as security, compliance, and operations.
  • Led requirement discovery sessions with stakeholders to translate regulatory, operational, and security requirements into explainable AI-driven risk assessment workflows.
  • Built a KPI-based risk scoring engine that normalizes multi-domain metrics and computes weighted risk scores, enabling consistent and explainable risk evaluation.
  • Implemented an automated reporting system generating structured risk reports with domain-wise breakdowns, key findings, and actionable recommendations- giving enterprise stakeholders a consistent and auditable basis for high-stakes technology decisions.
  • Established enterprise AI governance frameworks, architecture standards, deployment guidelines, reusable component libraries, and evaluation frameworks adopted across multiple delivery teams, reducing development redundancy by 40% and improving team productivity by 25%.
  • Designed automated LLM evaluation pipelines and AI governance workflows to assess enterprise AI solutions across quality, security, compliance, operational readiness, and production deployment, significantly reducing manual evaluation effort.
  • Influenced enterprise AI strategy by participating in technology evaluations, AI roadmap planning, and prioritization of AI initiatives based on business value, implementation complexity, organizational readiness, and expected impact.
  • Supported organization-wide AI adoption through reusable accelerators, architecture reviews, solution evaluations, technical mentoring, and governance best practices.
  • Reviewed enterprise AI solution proposals to validate technology selection, implementation feasibility, governance alignment, security, scalability, and long-term maintainability before project execution.
  • Supported development of two types of chatbots: (1) legacy FAQ-based bots using SQL-backed knowledge; (2) modern RAG-based intelligent bots tailored for different customers and datasets.
  • Vector DB created using Qdrant and knowledge graph using n8n for intent classification.
  • Member of the technical hiring team for new hires, interviewed AI, ML, NLP, and software engineering candidates across multiple experience levels, evaluating technical depth, architectural thinking, problem-solving ability, and delivery readiness.
A

Data Engineer, Data Science team

Argus India Price Reporting LLP. (India) · Jan 2024 – Apr 2025

Not yet confirmed
  • Built an AI-driven text generation solution leveraging NLP, Gemini LLM models, LangChain, AWS S3, OpenSearch and Generative AI techniques, with active editorial workflows supporting 145,000+ annual articles and 800 daily market-moving data stories, to extract key insights from historical text data of articles stored externally in a database using RAG and track trends based on keyword analysis which is integrated into a R Shiny app through APIs and deployed on posit connect.
  • Collaborated with editorial leadership, business stakeholders and global data teams to prioritize product priorities and enhancements, define implementation roadmap, and deliver AI-powered capabilities aligned with business workflows.
  • Processes for gathering data, calculating prices and other related work structures by writing strong and automated formulas and methodologies in R and producing reports and data feeds.
  • Perform proper data validation and quality assurance of every price and related process, formulas, and structures.
  • Migrated 4 legacy pricing pipelines from Excel-Access macro chains to Oracle and R Shiny - eliminating 1-6 hours of manual processing per pricing update and replacing fragile cross-file dependencies with scalable, auditable, real-time dashboards directly feeding data across 12 commodity categories including crude oil, LPG, coal, agriculture, chemicals, and metals, improving automation, data quality, and maintainability across processes and products.
  • Developed production-grade, scalable R Shiny dashboards and reusable R packages using the Golem architecture enforcing modular, standardized design patterns across all applications supporting 40,000+ live price data points published globally followed best practices for debugging, unit testing, and quality checks, and deployed apps using Posit Connect.
  • Performed data cleansing, data quality checking and management of data for product development processes, built predictive analytics and real-time time-series forecasting models using regression and statistical modeling to project future price scenarios across 12+ commodity categories, enhancing interpretability and supporting strategic pricing decisions on prices data enabling a 600+ people strong global editorial team to move from reactive reporting to proactive, data-driven market intelligence.
  • Interactive data visualizations for possibility curves along with confidence levels the model has at different prediction levels, with percentage-based confidence scores- translating complex statistical outputs into intuitive decision-support tools.
  • Designed and implemented ML-based anomaly detection models to monitor mismatch of records between development, testing and production environments and ensure the integrity of price trends, users, traders, deals and other records reducing manual oversight and integrated with dashboards.
  • Support clients with queries relating to integrating Argus data and metadata into client systems.
  • Prepare EDA tables and interactive plots with the help of plotly etc. to present to internal editorial teams and to internal and external clients.
  • Using project management tools, for example Jira, to manage all projects including new price creation, price creation, price codes, and many others.
Argus Media

Data Engineer

Argus Media · Jan 2024 – Apr 2025

Not yet confirmed
T

Business Process Lead, Statistical Monitoring Team

Tata Consultancy Services (TCS) · Nov 2022 – Jan 2024

Not yet confirmed
  • Led Statistical Monitoring and QTL teams delivering fraud detection and data integrity frameworks for one of the world's largest pharmaceutical companies across 2,000+ active clinical studies spanning 35,000 sites and medical centers globally, with findings directly informing regulatory submissions and patient safety decisions.
  • Coordinated cross-functional delivery across statisticians, developers, client stakeholders and clinical domain experts to ensure timely delivery of regulatory analytics solutions.
  • Finalized Statistical Monitoring Plans and QTL Plans analyze the clinical trials data, translating clinical trial requirements into the governing analytical blueprints used across all downstream fraud detection and quality monitoring activities and assign and support team members to perform the quality tailoring check to detect crucial and potential frauds using SAS, R, R Shiny and JMP Clinical.
  • Manage Jira, Trello for run assignment and runs tracking ServiceNow and Atlassian for study-wise data management, collaborate with client study team for information gathering, plan discussion, timelines, report discussions, etc.
  • Developed JMP-equivalent R Shiny dashboards for statistical tests complete with modals, accuracy measurement, logging, and automated report generation - replacing legacy tooling with modern, maintainable applications scaling across 35,000 global sites and 15-20 member study teams.
  • Designed and executed statistical fraud detection frameworks applying various statistical methods including demographic distribution, birthdate test, cluster analysis, perfect schedule of attendance, study visit, constant findings, duplicates records, multivariate outlier and inliers and adverse event summary with the help of JMP software and R Shiny applications, flagging statistically significant anomalies across 2,000+ studies, with findings escalated through regulatory compliance process.
  • Worked on integrating a natural language interface into R Shiny dashboards using LLMs, enabling users to query clinical trial data conversationally and receive dynamic visualizations based on statistical test results.
  • Based on the analysis, perform standard analysis to identify unusual or clustered patterns that have the potential to impact the data integrity of the study and/or may potentially impact patient safety and prepare standard report files.
  • Made two R packages following golem architecture using Bayesian methodology and Monte Carlo Simulation methods to compare posterior data with prior data at specified intervals of clinical studies to uncover data fabrication, perform fraud detection and related analysis, creating Excel and HTML files using R-markdown scripts for adverse events (AE), SAE, BPD and OPD (Protocol Deviations) and multi-site visualization plots - providing standardized, audit-ready comparative reporting across global clinical sites.
  • Wrote unit tests using devtools, shinytest2 and testthat package and automated delpoyment with GitHub Actions.
  • Developed and validated ML-based models for pattern recognition in adverse event reporting, contributing to regulatory compliance and clinical safety insights and deployed on Posit Connect.
  • To prepare methodologies, related documentation, and do the modeling on SDTM datasets and convert the results to reports to process further for quality check.
  • Provide useful suggestions to improve and automate current processes and development of new versions of current package, evaluation, testing and validation to ensure quality and standardization.
  • Managed Table Listing Figures (TLFs) to organize and present data, automating reports with R to deliver accurate, timely data insights.
Tata Consultancy Services

Statistical Programmer and Data Analyst

Tata Consultancy Services · Nov 2022 – Jan 2024

Not yet confirmed
I

Academic Associate, Operations Management and Quantitative techniques (OM & QT)

Indian Institute of Management, Indore (M.P.) · Nov 2019 – Nov 2022

Not yet confirmed
  • To prepare standard research methodologies, decision-making techniques, and statistical methods for scraping and analyzing data to prepare research reports, case writings, and academic publications.
  • Conducted research and implemented machine learning algorithms such as classification, regression, clustering, time- series forecasting, natural language processing analysis, and survival models to support decision-making in academic studies.
  • Explored and adapted deep learning architectures and optimization strategies based on research to improve the performance of analytical solutions in operations research projects.
  • Extracted and preprocessed text from PDF-based research papers using web scraping, OCR, and NLP pipelines, and applied TF-IDF and frequency analysis to identify dominant terms within topic-specific corpora.
  • Built a logistic regression model on student admission data to predict enrollment likelihood, identifying key features influencing decision outcomes.
  • Applied a range of machine learning algorithms including logistic and linear regression, decision trees, random forest, XGBoost, SVM, KNN, PCA, and clustering techniques for predictive modeling, dimensionality reduction, and pattern recognition across diverse datasets.
Indian Institute of Management, Indore

Academic Associate

Indian Institute of Management, Indore · Nov 2019 – Nov 2022

Not yet confirmed

Skills 0 proven through work

Also works with

Azure Document Services UsageMentoring and Peer CoachingTeamworkAdaptabilityProbability AnalysisData StorytellingClient FocusWorking with Real-world DatasetsData Integrity ManagementFraud Detection System DevelopmentDatabase ManagementRetrieval-Augmented GenerationDocument ChunkingLarge Dataset HandlingAWS DeploymentData PreprocessingETL Pipeline DevelopmentData ExtractionAutomated Summarization SystemsNatural Language GenerationExplaining Technical ConceptsTechnology Stack SelectionTechnical Feasibility AssessmentCost EstimationUse Case DevelopmentFeasibility and Impact AssessmentWork Planning and PrioritizationRequirements AnalysisProduct DevelopmentProject Architecture Planning

Proof of Work

Proof of Work

Surbhi shares this with people who ask. You'll hear back either way.

Education

Academic Associate, Operations Management and Quantitative techniques (OM & QT)

Indian Institute of Management, Indore (M.P.) · 2019 — 2022

Master of Science, Statistics

School of Statistics, Devi Ahilya Vishwavidyalaya, Indore · 2017 — 2019

Bachelor of Science, Computer Science, Statistics & Mathematics

Govt. Holkar Science College, Indore (M.P.) · 2014 — 2017

Contact details

Contact details

Surbhi shares this with people who ask. You'll hear back either way.

Ask about this candidate

AI responses may contain errors. Proof status and source information are provided by Proof of Skill.

A Proof CV — one profile, kept current, with the proof attached. Back to the portal