Kaushik

Kaushik

Open to

Bangalore · Mumbai · Hyderabad

Work style

remote, hybrid, onsite

Contact details — on requestProof of Work — on request

Experience

I

AI Data Engineer

Intellirev · May 2024 – Jul 2026

Not yet confirmed
  • Built real-time ETL pipelines in Python (Pandas, Scikit-learn) orchestrated via Cloud Functions and Cloud Scheduler, achieving 99.9% data uptime while processing $1M+ in annual claims and payments for production AI agent deployment.
  • Built data quality checks, covering schema validation and anomaly thresholds across the claims pipeline, catching bad records and cutting downstream data errors by 25%.
  • Integrated Machine Learning models and LLM-based agents into ELT workflows, cutting manual reconciliation time by 40% and flagging 95% of billing anomalies before submission.
  • Migrated the claims pipeline to a fully serverless, HIPAA-compliant GCP architecture (Cloud Storage, BigQuery), handling 500+ monthly transactions and eliminating third-party infrastructure licensing overhead.
  • Partnered with data scientists and BI stakeholders to operationalize ML-driven claim analysis via Datastream and BigQuery pipelines, increasing collected revenue 15%.
I

AI Data Engineer

Intellirev · May 2024 – Jun 2026

Not yet confirmed
  • Built real-time ETL pipelines in Python (Pandas, Scikit-learn), orchestrated via Cloud Functions and Cloud Scheduler, holding 99.9% data uptime while processing $1M+ in annual claims and payments for production AI agent deployment.
  • Designed schema validation and anomaly-threshold checks across the claims pipeline, cutting downstream data errors by 25%.
  • Integrated ML models and LLM-based agents directly into ETL workflows, cutting manual reconciliation time by 40% and flagging 95% of billing anomalies before submission.
  • Led the migration of the claims pipeline to a fully serverless, HIPAA-compliant GCP architecture (Cloud Storage, BigQuery), eliminating third-party licensing overhead while scaling to 500+ monthly transactions.
  • Partnered with data scientists and BI stakeholders to operationalize ML-driven claim analysis via Datastream and BigQuery pipelines, driving a 15% increase in collected revenue.
BNY Mellon

Data Engineer

BNY Mellon · Jul 2023 – Mar 2024

Not yet confirmed
  • Engineered an end-to-end fraud detection pipeline in Python (pandas, scikit-learn), applying SMOTE-based class rebalancing to improve fraud recall from 65% to 85% on 250,000+ transaction records, achieving 97.4% ROC-AUC with Random Forest
  • Built and maintained ETL pipelines on Databricks, processing records with PySpark and Databricks notebooks, reducing query times by 20-40% through pipeline optimization and workflow automation
  • Partnered with AI engineering and audit teams to define data-governance protocols using a data catalog and validation rules, improving overall data-quality compliance
  • Designed and optimized SQL data pipelines and schemas, writing complex queries (joins, CTEs, aggregations) to support reporting and analytics workflows, improving efficiency by 30%
  • Trained and integrated ML models on Databricks, combining Spark-based feature pipelines with scikit-learn and native MLlib, enabling real-time analytics for downstream applications
BNY Mellon

Data Engineer

BNY Mellon · Jul 2023 – Feb 2024

Not yet confirmed
  • Engineered an end-to-end fraud detection pipeline in Python (pandas, scikit-learn) on 250,000+ transaction records, applying SMOTE-based class rebalancing to lift fraud recall from 65% to 85%, achieving 97.4% ROC-AUC with Random Forest.
  • Built and maintained ETL pipelines on Databricks with PySpark, cutting query times 20-40% through pipeline optimization and workflow automation.
  • Defined data-governance protocols with AI engineering and audit teams using a data catalog and validation rules, improving data-quality compliance across the pipeline.
  • Optimized SQL data pipelines and schemas, writing complex joins, CTEs, and aggregations to support reporting workflows, lifting reporting efficiency 30%.
  • Combined Spark-based feature pipelines with scikit-learn and native MLlib on Databricks, enabling real-time analytics for downstream applications.
A

Data Analyst

Awaywegoo · Jul 2019 – Jul 2021

Not yet confirmed
  • Engineered ETL pipelines and complex SQL transformations within MySQL, increasing data throughput by 40% and cutting BI reporting latency by over 2 days to accelerate executive decision-making.
  • Built reusable dashboard templates for investment teams, driving adoption of self-service analytics and freeing analysts from repetitive report requests.
  • Architected an enterprise PostgreSQL analytics platform that executes 10K+ monthly queries and reduces cycle times for Go-To-Market (GTM) and supply chain strategies.
AwayWeGoo

Data Analyst

AwayWeGoo · Jun 2019 – Jul 2021

Not yet confirmed
  • Engineered ETL pipelines and complex SQL transformations in MySQL, increasing data throughput by 40% and cutting BI reporting latency by 2+ days.
  • Built reusable dashboard templates for investment teams, driving adoption of self-service analytics and freeing analysts from repetitive report requests.
  • Architected an enterprise PostgreSQL analytics platform running 10K+ monthly queries, cutting cycle times for go-to-market and supply chain strategy work.

Skills 0 proven through work

Also works with

Common Table ExpressionsDelta Engine Features UsageWork Planning and PrioritizationTeamworkProcessing Pipeline OptimizationJupyter NotebookData LoadingData ExtractionData-Driven RecommendationsData StorytellingTabular Data Cleaning & ValidationLow-Latency System OptimizationCost-Saving Initiative ManagementGenerative AI ApplicationData Architecture DesignDatabase Index CreationRoot Cause AnalysisDatabase Query OptimizationPostgreSQL Database UsageTechnical TroubleshootingReporting Time ReductionBusiness Growth StrategyData TransformationServerless ArchitectureStreaming Data ProcessingData GovernanceData OptimizationData CatalogingDelta Live Tables UsageLegacy System MigrationPySparkDatabricks UsageSMOTE UsageHandling Data ImbalanceAnomaly DetectionGreat Expectations UsagePydantic UsageFraud Detection System DevelopmentWorkflow CoordinationData ValidationDocument ParsingLLM Application IntegrationPython DevelopmentETL Pipeline DevelopmentData Engineering

Proof of Work

Proof of Work

Kaushik shares this with people who ask. You'll hear back either way.

Education

Master's, Information Systems

Stevens Institute of Technology · 2021 — 2023

Bachelor's, Computer Engineering

Mumbai University · 2016 — 2019

Contact details

Contact details

Kaushik shares this with people who ask. You'll hear back either way.

What drives their work

The Architect — Builds what holds

The Architect

Designs and refines the underlying structure for robust data flow.

Kaushik's superpower

You excel at diagnosing data infrastructure weaknesses and designing scalable, high-performance solutions.

How Kaushik works

You would excel in roles focused on designing, building, and optimizing robust data platforms and pipelines.

Ask about this candidate

AI responses may contain errors. Proof status and source information are provided by Proof of Skill.

A Proof CV — one profile, kept current, with the proof attached. Back to the portal