Skip to main content
VHire
Back to Jobs
N

Data Engineer

NITYOMumbai, Maharashtra, IndiaPosted 2 weeks ago• Updated 1 weeks ago
Contract
Mid level
On-site

Salary

5 - 12 LPA

Experience

5-9 Years

VHire page views

42

VHire applicants

0

Activity recorded on VHire only. Page views are not unique visitors; applicants from employer sites and ATS systems are not included.

Data Engineer

This Data Engineer position involves building and managing data pipelines, vector infrastructure, and real-time processing architectures. The ideal candidate will have experience with distributed data systems, SQL, and orchestration tools.

Key Responsibilities

  • Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency.
  • Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers.
  • Implement zero-trust access, privacy controls, and compliance within AI context pipelines.
  • Build event-driven architectures for real-time processing and continuously refresh embeddings and indexes.

Requirements

  • Deep experience with distributed data systems, SQL, and orchestration tools.
  • Experience tuning high-throughput database infrastructure.
  • Knowledge of Google's GECX is a plus.
  • Familiarity with chunking strategies and embedding models.

Work Arrangements

  • Location: Mumbai, Maharashtra, India
  • Type: Contract
  • Remote: Not specified

Skills

  • ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency.
  • Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization.
  • Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures.
  • Data Governance: Zero-trust access, privacy controls, compliance, and auditability.
  • Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems.
  • Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection.
  • Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms.
  • Performance Tuning: High-throughput read/write optimization.
  • Data Quality & Lineage: Validation, schema enforcement, and lineage tracking.

Your application is completed on VHire. Review the role requirements before starting.

Required Skills
ETL & Data Modeling
Search Infrastructure
Performance Tuning
Real-time Processing
SQL
Chunking & Embeddings
Data Quality & Lineage
Orchestration Tools
Vector Databases
Distributed Data Systems
Data Engineer
Data Governance
Azure AI Search
pgvector
Inverted Indexes
Kafka
Ranking Algorithms
Great Expectations
Redis
Flink
BM25
Google's GECX
OpenLineage
Spark