Associate Director Product
Bangalore, KA, IN, 560100
Summary
This role is strictly involved in business analysis, requirements documentation, and healthcare reference data support activities and does not involve direct access to Protected Health Information (PHI), Personally Identifiable Information (PII), or any secured or confidential client data. The work is limited to analysis of healthcare claims, reference code sets, and system configurations using governed datasets and does not include handling or processing of sensitive health or personal information.
We’re seeking a hands-on engineer who blends GenAI expertise with solid data engineering skills to design, build, and operationalize AI-enabled solutions on Azure. You’ll work across the lifecycle—from ingesting and modeling data to orchestrating LLM-powered workflows (incl. RAG)—with strong emphasis on reliability, security, cost efficiency, and measurable outcomes. The ideal candidate is comfortable converting moderately detailed requirements into robust designs and high-quality code.
Your role in our mission
LLM Solution Engineering
Design and implement retrieval-augmented generation (RAG) pipelines: document preprocessing, chunking, embeddings, hybrid retrieval, grounding/citations, prompt templating.
Build and integrate MCP (Model Context Protocol) tools/servers to expose enterprise systems and data sources to LLMs securely.
Develop orchestration layers (tool/function calling, guardrails, prompt contracts) using Azure OpenAI with frameworks like Llamaindex.
Build custom tools that can seamlessly interact with GitHub Copilot to act as force multiplier for the developers.
Data Engineering on Azure
Build ingestion and transformation pipelines into ADLS Gen2 (bronze/silver/gold), using efficient formats (Parquet, optionally Delta).
Implement ETL/ELT workflows and scheduling (e.g., Data Factory, Functions, Container Apps), with schema design for telemetry, operational metrics, and curated aggregates.
Optimize storage layout (partitioning, file sizing), handle schema evolution, and maintain data quality checks and lineage.
Analytics & Experimentation
Operationalize statistical analyses (joins, aggregations, time-series); run A/B tests, t‑tests, and basic causal methods (e.g., difference‑in‑differences) to quantify impact.
Deliver consumable datasets with clear KPIs and semantic modeling; instrument cost/latency/quality metrics for AI features.
Quality, Security & Governance
Integrate quality signals from external reports; correlate them with engineering and product metrics for insights.
Implement security best practices: data boundary controls, role-based access, secrets management (Key Vault), tenant isolation, audit logging; align with responsible AI guidelines.
Platform & DevOps
Productionize services on Azure (Functions, App Service/Container Apps, AKS as needed); implement CI/CD (GitHub Actions/Azure DevOps), testing, observability (App Insights/Log Analytics).
Monitor performance and cost; apply model selection, token optimization, caching, and batching strategies.
Design & Collaboration
Translate moderately specified requirements into architecture docs, ADRs, and sequence diagrams; drive cross-functional alignment with product, engineering, and QA.
What we're looking for
Experience (5–7 years) delivering production systems that combine data engineering with LLM applications.
Azure OpenAI hands-on experience (prompting, model selection, evals, guardrails) and LLM orchestration using Semantic Kernel or LangChain.
RAG expertise: embeddings, hybrid search, chunking, query rewriting, grounding with citations; Azure AI Search (vector + keyword).
MCP (Model Context Protocol): building/integrating tools/servers to expose enterprise data to models securely.
Data Engineering (Azure): ADLS Gen2 (bronze/silver/gold), ETL/ELT pipelines (Data Factory/Functions/Containers), Parquet; solid schema design and partitioning strategies.
Programming: Strong Python (pandas), solid SQL; working TypeScript for integration or UI layers.
Analytics: Practical statistics (A/B tests, t‑tests, DiD), time-series analysis; Power BI modeling and dashboards.
DevOps: CI/CD (GitHub Actions/Azure DevOps), Docker; secrets and configuration management (Key Vault); logging/monitoring (App Insights).
Security & Compliance: Role-based access, data boundaries, PII/PHI awareness, audit trails; responsible AI practices.
Knowledge graphs/traceability modeling; semantic caching strategies.
Evals frameworks (e.g., Giskard, DeepEval) and custom evaluation harnesses for groundedness/accuracy/cost.
Azure OpenAI integration and prompt engineering for enterprise use cases
Document Intelligence (Azure Form Recognizer, OCR pipelines) and processing unstructured data using Python libraries (e.g., Pandas, PyMuPDF, Tesseract)
Performance optimization of RAG (Retrieval-Augmented Generation) solutions, including latency reduction and cost efficiency.
Advanced indexing mechanisms for vector databases (e.g., HNSW, IVF, PQ) and hybrid search strategies.
What you should expect in this role
Reliable, secure, and cost-efficient AI-enabled features backed by measurable improvements.
Maintainable data estate with clear lineage and curated datasets for decision-making.
Reusable components and patterns that accelerate future AI/data initiatives across the organization