SS

siri sampelli

As a Data Engineer with two years of experience, I am driven by a meticulous and diligent approach to building robust data solutions. My ambition for mastery ensures I consistently deliver high-quality, accountable results, always striving for precision and logical efficiency in my work.

Open to

Hyderabad · Bangalore

Work style

on-site

Contact details — on requestProof of Work — on request

Experience

B

Data Engineer II

BANK OF AMERICA · Jul 2019 – Present

Not yet confirmed
  • Senior Data Engineer with 7+ years of experience building and supporting scalable data pipelines and distributed data processing solutions in enterprise environments.
  • Expertise in Python, SQL, PySpark, Spark SQL, Hive, with strong experience in data modeling, schema design, data quality, reconciliation, and performance optimization.
  • Skilled in working with large-scale datasets using JSON, Avro, ORC and Parquet formats, and implementing reliable ETL solutions using CI/CD and software engineering best practices.
  • Proven track record of delivering production-ready data solutions, leading technical initiatives, and collaborating with cross-functional teams in Agile environments.
  • Designed and developed PySpark batch pipelines to ingest and process large-scale datasets from Oracle, HR systems, and access management platforms, storing curated data in partitioned Parquet/ORC Hive tables for IDV compliance reporting and analytics.
  • Built ingestion framework to load Aadhaar and Passport based verification records into Hive raw layer, and developed Spark transformations to validate IDV status against associate data, privileged user lists, and HR employee attributes.
  • Implemented business rules to identify non-compliant users by validating IDV completion, HYPR enhanced authentication flag, and exception accounts sourced from Splunk logs.
  • Developed curated and consumption layer Hive tables containing IDV compliance metrics such as verified users, pending verification, privileged users without IDV, and exception accounts.
  • Implemented comprehensive data quality and reconciliation controls including record count validation, null checks, duplicate detection, schema validation, and cross-source consistency checks to ensure accuracy and completeness of compliance datasets.
  • Published IDV compliance datasets to Tableau dashboards and supported audit teams by delivering actionable metrics.
  • Developed curated, high-quality compliance datasets with validation and reconciliation controls, enabling downstream analytics and AI-driven reporting use cases across identity governance and compliance functions.
  • Consolidated Identity and Access Management data including users, roles, entitlements, and service accounts from multiple source systems, and built PySpark transformations to standardize and integrate datasets for governance reporting.
  • Implemented service account attestation logic by identifying service accounts, mapping owners, and generating certification status to support periodic access reviews and compliance requirements.
  • Developed dormancy checks using last login and activity timestamps to detect inactive user accounts and unused service accounts, enabling governance teams to take remediation actions.
  • Built logic to identify privileged and elevated access across critical applications by analyzing role-to-entitlement mappings and flagging high-risk access scenarios.
  • Designed curated Hive tables containing aggregated IAM metrics including dormant accounts, service account attestation status, privileged access, and orphan accounts for downstream reporting.
  • Published IAM governance datasets to Tableau dashboards and supported audit teams by delivering compliance metrics and resolving data discrepancies.
Bank of America

Data Engineer 2

Bank of America · Jul 2019 – Present

Not yet confirmed
  • Senior Data Engineer with 6 years of experience building scalable ETL pipelines in banking environments, processing large-scale datasets for compliance and analytics.
  • Expertise in Python, PySpark, Spark SQL, and Hive, with a strong focus on data validation, reconciliation, and performance tuning.
  • Delivered high-quality, audit-ready datasets and optimized pipeline performance in production environments.
  • Proven track record of developing reliable data solutions within Agile teams using CI/CD practices

Skills 0 proven through work

Also works with

Governance ReportingWorking with CSV FilesProactive InitiativeProblem SolvingConflict Resolution in TeamsTransparent CommunicationStakeholder Communication AdaptationTeamworkFull-stack DevelopmentWork Planning and PrioritizationData Architecture DesignTrade-off AnalysisTechnical Solution DesignPython DevelopmentAWS DeploymentPySparkData OptimizationData Partitioning and Segregation DesignAnalytical Report CreationApplication DeploymentAWS Secrets Manager UsageRule-Based System DesignData Repetition PreventionData PreprocessingData ValidationTask SchedulingApache KafkaStreaming Data ProcessingBatch Processing ImplementationLow-Latency System OptimizationData Quantity ManagementRequirements AnalysisData TransformationTabular Data Cleaning & ValidationData CollectionDatabase IntegrationAPI IntegrationData LoadingETL Pipeline Development

Proof of Work

Proof of Work

siri shares this with people who ask. You'll hear back either way.

Education

B.Sc., Computer Science

G. Narayanamma Institute of Technology and Science - Hyderabad · 2015 — 2019

Intermediate, Intermediate

Sri Vikas Junior College, Warangal · 2013 — 2015

High School, 10th Board

Daffodil's high school, Warangal · 2012 — 2013

Contact details

Contact details

siri shares this with people who ask. You'll hear back either way.

Ask about this candidate

AI responses may contain errors. Proof status and source information are provided by Proof of Skill.

A Proof CV — one profile, kept current, with the proof attached. Back to the portal