I am an AI & Data Engineer at Google based in Hyderabad. I specialize in architecting distributed data lakehouses and production agentic AI systems for major banks and enterprise financial workloads.
I am an AI & Data Engineer at Google with over 10+ years of global experience architecting mission-critical data platforms and distributed systems.
My engineering focuses on building high-throughput Big Data lakehouses and operationalizing production-grade Agentic AI workflows—partnering with major banks and global financial institutions to deliver scalable, secure, and deterministic architectures.
A selection of production lakehouses, GraphRAG engines, and open Model Context Protocol tooling.
Slide (8)
AGENTIC AIOpen Source
Equity Portfolio Analyzer Agent (Google ADK)
Multi-source financial agent built on Google ADK fanning out across 4 market data feeds (Yahoo Finance, Twelve Data, NSE/BSE, Tavily) to synthesize an opinionated BUY / HOLD / SELL verdict.
4 Parallel
Data Feeds
~115 Metrics
Data Points
Google ADK
Framework
Google ADKLiteLLMPythonFastAPIFinancial APIsNSE/BSE
MCP TOOLINGOpen Source
Production Database Model Context Protocol (MCP) Suite
Enterprise MCP servers enabling AI agents to safely introspect schemas, generate optimized SQL, and execute database tasks with strict guardrails.
< 140ms
Invocation Latency
100%
Safety Pass Rate
Cloud Run / Docker
Supported Runtimes
Model Context Protocol (MCP)FastAPIGoogle Cloud RunBigQueryCloud SQLDocker
DATA ENGINEERINGOpen Source
GCP Anti-Money Laundering (AML) AI Framework
Enterprise Anti-Money Laundering (AML) transaction compliance pipeline and dashboard deployed on GCP using FastAPI, BigQuery, and Cloud SQL.
Banking AML
Target Domain
BigQuery / SQL
Query Engine
FastAPI / GCP
Backend
Google CloudBigQueryFastAPICloud SQLPythonSQL DDL/DML
AGENTIC AIOpen Source
Agentic AI Observability & Telemetry Framework
Distributed tracing and evaluation framework tracking multi-agent reasoning steps, tool latencies, and token cost telemetry into GCP Cloud Trace.
< 4ms
Trace Overhead
Near Real-Time
Eval Pipeline SLA
Per-Agent Granular
Cost Attribution
OpenTelemetryGoogle Cloud TracePythonVertex AI PipelinesPrometheus
AGENTIC AIOpen Source
Automated Knowledge Graph Construction via LLM
Pipeline leveraging few-shot prompt chaining and graph reconciliation to convert unstructured domain documents into queryable property graphs.
Containerized, auto-scaling API services on Google Cloud Run delivering low-latency semantic search and analytical agent endpoints.
< 800ms
Cold Start
99.99%
Availability
Thousands req/min
Throughput
FastAPIGoogle Cloud RunDockerCloud SQLPythonGCP
Work Experience
A decade of engineering distributed data systems, enterprise lakehouses, and production Agentic AI across Google, Carelon, Deloitte, and Cognizant.
AI & Data Engineer
Google
Hyderabad, India•2022 – Present
Current Role • Featured Impact
Partnering with major banks, financial institutions, and global enterprises to architect mission-critical Big Data lakehouses and deploy production Agentic AI systems with strict security guardrails.
Key Engineering Deliverables & Impact at Google
Engineered production agentic AI workflows and Model Context Protocol (MCP) database toolboxes for financial workloads with AST-level query validation.
Designed sub-second GraphRAG retrieval pipelines combining Cloud Spanner Graph and BigQuery for complex multi-hop financial knowledge reasoning.
Implemented high-volume CDC streaming and lakehouse physical models processing 50k+ events/sec benchmarked against enterprise TPC-DS workloads.
Pioneered end-to-end LLM agent observability using OpenTelemetry to capture reasoning traces, latency, token economics, and evaluation scores.
Core Stack:BigQueryCloud SpannerPySparkLangGraphModel Context ProtocolVertex AIOpenTelemetryCloud Run
Previous Enterprise Experience
Select a company to view role details
Senior Data Engineer
Carelon
Engineered scalable Big Data pipelines and cloud analytics infrastructure for massive healthcare and enterprise workloads.
2021 – 2022Hyderabad, India
Built distributed PySpark pipelines processing multi-TB transactional healthcare and claims records.
Engineered automated ETL workflows and validation frameworks ensuring zero data drift.
Optimized query runtimes and partitioning models on distributed storage for executive reporting.
Collaborated across international engineering teams on production cloud data platforms.
A comprehensive guide to capturing distributed spans, metrics, and agent reasoning traces in Google Agent Development Kit (ADK) using OpenTelemetry and Cloud Trace.
Architecting high-throughput, cost-effective data migrations from on-premise Apache Hive metastores to Google BigQuery using serverless PySpark on Dataproc.
Google BigQueryDataproc ServerlessApache HivePySpark