You're seeing this page as if you were . The main menu is still yours, though. Exit from immersion
Vishal AwateVA

Vishal Awate

LEAD DATA ENGINEER · DATA PLATFORMS · GenAI / LLM

600 €/jour
Pune, IN
8-15 ans

Délai de réponse moyen : 1h

À propos de Vishal

Lead Data Engineer with 9+ years of experience designing petabyte-scale data platforms — and a hands-on builder of production GenAI/LLM systems.

I build GenAI for a global banking client — a production service that translates complex SQL and multi-language code (Python, Scala, Java, and more) into plain-English business explanations: hybrid RAG combining deterministic structured retrieval with vector search, database-driven prompt management, LLM-as-judge quality evaluation, and fully deterministic outputs built for regulated environments, running on Llama 3.3 through an enterprise LLM gateway.

Independently, I designed and operate a second production GenAI system end to end — multi-source data ingestion (25+ sources), an LLM pipeline with structured extraction using tool schemas, embedding-based semantic search (RAG) with vector storage in PostgreSQL, multi-stage match scoring with model cascading that keeps inference costs under control, LLM-as-judge evaluation gating notifications, and LLM-generated documents delivered over WhatsApp and email.

I architect end-to-end data solutions: petabyte-scale data lake pipelines in the payments domain (PayPal/Braintree) using Spark/PySpark, lakehouse platforms on Azure Databricks, and the orchestration, quality, and cost controls that keep them reliable in production — plus platform governance frameworks automating storage, compliance, and cost monitoring across 10+ Hadoop databases serving 15+ teams.


Tech: Spark · PySpark · Python · Azure Data Factory · Azure Databricks · AWS (S3, Redshift, Lambda) · GCP core services · Hive · Impala · HDFS · SQL (MySQL, Oracle, PostgreSQL) · Shell · CI/CD (Jenkins, uDeploy)

AI: LLMs (Claude, OpenAI, Llama 3.3) · Hybrid RAG (structured + vector) · Embeddings & Vector Search (pgvector) · AI Agents · Prompt Engineering · LLM-as-Judge · Model Cost Optimization · Deterministic LLM Outputs
  • Anglais

    Bilingue ou natif

En télétravail uniquement
Travaille majoritairement à distance

Expériences

  • Atyeti Inc
    Lead Associate — Data Engineering Lead
    mai 2024 - Aujourd'hui (2 ans et 3 mois)
    Pune, Maharashtra, India
    Lead Data Engineer with 9+ years of experience designing petabyte-scale data platforms — and a hands-on builder of production GenAI/LLM systems.

    I build GenAI for a global banking client — a production service that translates complex SQL and multi-language code (Python, Scala, Java, and more) into plain-English business explanations: hybrid RAG combining deterministic structured retrieval with vector search, database-driven prompt management, LLM-as-judge quality evaluation, and fully deterministic outputs built for regulated environments, running on Llama 3.3 through an enterprise LLM gateway.

    Independently, I designed and operate a second production GenAI system end to end — multi-source data ingestion (25+ sources), an LLM pipeline with structured extraction using tool schemas, embedding-based semantic search (RAG) with vector storage in PostgreSQL, multi-stage match scoring with model cascading that keeps inference costs under control, LLM-as-judge evaluation gating notifications, and LLM-generated documents delivered over WhatsApp and email.

    I architect end-to-end data solutions: petabyte-scale data lake pipelines in the payments domain (PayPal/Braintree) using Spark/PySpark, lakehouse platforms on Azure Databricks, and the orchestration, quality, and cost controls that keep them reliable in production — plus platform governance frameworks automating storage, compliance, and cost monitoring across 10+ Hadoop databases serving 15+ teams.


    Tech: Spark · PySpark · Python · Azure Data Factory · Azure Databricks · AWS (S3, Redshift, Lambda) · GCP core services · Hive · Impala · HDFS · SQL (MySQL, Oracle, PostgreSQL) · Shell · CI/CD (Jenkins, uDeploy)

    AI: LLMs (Claude, OpenAI, Llama 3.3) · Hybrid RAG (structured + vector) · Embeddings & Vector Search (pgvector) · AI Agents · Prompt Engineering · LLM-as-Judge · Model Cost Optimization · Deterministic LLM Outputs
    LLM Apache Spark PostgreSQL RAG AI Agent
  • Clairvoyant LLC (an EXL Company)
    Senior Software Engineer
    décembre 2017 - avril 2024 (6 ans et 4 mois)
    Pune, Maharashtra, India
    • • Core engineer on PayPal/Braintree's enterprise data lake (CPDL) — petabyte-scale payments data across ingestion, processing, and data-quality layers (Braintree, Platform, and Swift workstreams).
    • • Built a generic PySpark batch-ingestion framework with integrated data-quality checks (null, duplicate, trend analysis), adopted as the team standard for onboarding new hourly data sources — cutting source onboarding from custom builds to configuration.
    • • Developed AWS extraction pipelines (S3, EC2, Redshift, Boto) feeding the CPDL batch-ingestion framework, with UC4-scheduled daily production workflows — decryption and validation before ingestion into Hive, with transformations on top.
    • • Designed a CDM automation tool that mapped multiple source tables into Common Data Model tables automatically, eliminating repetitive manual modelling work.
    • • Optimized long-running Spark/Hive workloads via partitioning, bucketing, and file-size management on CPDL's multi-terabyte daily volumes.
    • • Delivered Lloyds Bank's mainframe-to-Hadoop migration, moving legacy banking data into HDFS/Hive with reconciliation checks.
    • • Led Azure cloud migration for a cybersecurity client (Abacus) — metadata-driven Azure Data Factory pipelines ingesting REST APIs, PostgreSQL, MySQL, and SQL Server into a medallion architecture (bronze/silver/gold) with Databricks notebook transformations, Key Vault-managed credentials, gold-layer data models, and Grafana monitoring.
  • Centurysoft Private Limited
    Hadoop Developer
    mai 2017 - octobre 2017 (5 mois)
    Pune, Maharashtra, India
    Social-Media Analytics Client
    • • First industry role in big data — built Hadoop-ecosystem data-processing jobs (HDFS, Hive, Sqoop, Pig), processing terabytes of daily log data for a social-media analytics client (Solvedrum) into controls and change-history reporting.

Recommandations

Soyez le premier à recommander Vishal

Contribuez à la réussite de ce freelance en partageant votre expérience de collaboration avec lui.

Ces profils de freelance correspondent également à vos critères

AgathaA

Agatha Frydrych

Backend Java Software Engineer

4.7

(3)

2

BaptisteB

Baptiste Duhen

Fullstack developer

4.6

(4)

5

AmedA

Amed Hamou

Senior Lead Developer

4

(2)

7

AudreyA

Audrey Champion

Web developer

4.3

(3)

4

Formations

  • B.Tech
    Shivaji University
    B.Tech

Compétences

Catégories