Data Engineer - ML & AI in Bengaluru, India - Jobeax
Vacancy description
Data Engineer - ML & AI in Bengaluru, India
IMerit Technology
India, Bengaluru
Data Engineer - ML & AI in Bengaluru, India is listed on Jobeax. Browse 30,000+ vacancies available.
iMerit is a leading AI data solutions company specializing in transforming unstructured data into structured intelligence for advanced machine learning and analytics applications. Our clients span autonomous mobility, medical AI, agriculture, and more—powering next-generation AI systems with high-quality data services. We are seeking a skilled Data Engineer to help scale and enhance our internal data observability and analytics platform. This platform integrates with data annotation tools and ML pipelines to provide visibility, insights, and automation across large-scale data operations. You will design and optimize robust data pipelines, build integrations with internal platforms (e.g., Design and build scalable batch and real-time data pipelines across structured and unstructured sources. Integrate analytics and observability services with upstream annotation tools and downstream ML validation systems to enable full-cycle traceability. Collaborate with product, platform, and analytics teams to define event models, metrics, and data contracts. Develop ETL/ELT workflows using tools like AWS Glue, PySpark, or Airflow; ensure data quality, lineage, and reconciliation. annotation throughput, quality KPIs, latency). Build data models and queries to power dashboards and insights via tools like Athena, QuickSight, or Redash. Contribute to infrastructure-as-code and CI/CD practices for deployment across cloud environments (preferably AWS). Document architecture, data flow, and support runbooks; continuously improve platform performance and resilience. Integrate with customer data platforms and pipelines, including bespoke data frameworks.
4–8 years of experience in data engineering or backend development in data-intensive environments. Proficient in Python and SQL; Strong experience with cloud-native data tools and services (S3, Lambda, Glue, Kinesis, Firehose, RDS). Experience with data lake and warehouse patterns (e.g., Delta Lake, Redshift, Snowflake). Solid understanding of data modeling, schema design, and versioned datasets. Data Governance and Security: Understanding and implementing data governanc policies and security measures. Proven experience in building resilient, production-grade pipelines and troubleshooting live systems. Good working knowledge of Database fundamentals, relational databases and SQL
Experience with observability/monitoring systems (e.g., Familiarity with data governance, RBAC, PII redaction, or compliance in analytics platforms. Exposure to annotation/ML workflow tools or ML model validation platforms. Comfort working in Agile, distributed teams using tools like Git, JIRA, and Slack.
You'll work at the intersection of AI, data infrastructure, and impact—contributing to platforms that ensure AI is explainable, auditable, and ethical at scale. Join a team building the next generation of intelligent data operations.
... expertise in NLP, Fundamental machine learning, deep learning, transformer, state space-based architecture.- Azure ML and/or AWS.- Strong in Python coding, SQL and database queries, data preparation, and analysis.- Exploratory Data Analysis (EDA).- Experience with PyTorch.- LLM training and fine-tuning (e.g., GPT, LLaMA, Mistral, ...
... Extensive experience in building, consuming, and optimizing RESTful APIs, with proficiency in tools like Swagger, Postman, or similar. - Strong knowledge of SQL databases and querying languages - Demonstrated experience in building and maintaining robust CI/CD pipelines using tools such as Jenkins or GitLab CI. - Exceptional ...
... the Fortune 50, achieve discoveries, insights, and business outcomes faster and more sustainably. We're passionate about solving our customers' most complex data challenges to accelerate intelligent innovation and business value. As a Senior/Staff Kernel Engineer at Weka, your primary responsibility will be collaborating ...
... Celery or other queue mechanism is Must Good to Have: - Experience working with Docker and cloud platforms (AWS/GCP/Azure) is good to have - Familiarity with ML model serving is Nice to have. - Exposure to GenAI (e.g., llama, OpenAI, HuggingFace, LangChain, Agentic AI etc) or Data Engineering (e.g., Data platforms, pandas, ...
Shift Timings: Design, build, and maintain robust ETL/ELT pipelines feeding a Snowflake-based data platform Build and manage integrations using SnapLogic to connect source systems, APIs, and downstream consumers Develop and maintain data models and transformations in dbt, including tests, documentation, and CI/CD-based ...
Job Title: Senior Data Engineer Eperience - 10+ to 18 yrs Location: Remote Type: Contract( Comfortable with a 6-month contractual role) Requires strong hands-E xperience in AI and Data Engineering with strong expertise in Azure OpenAI, Azure AI Foundry, Agentic AI, RAG, LLMs, AutoGen, OCR, Databricks, PySpark, Python, Azure ...
... performant Use AI tooling within DE workflows, including code generation, pipeline automation, and data quality checks Contribute to and help shape company-wide data governance standards Collaborate with analytics, BI, and business teams to deliver trusted, well-modeled data 5–7 years of hands-on data engineering experience ...
... Engineering, or related field. - 8+ years of experience in Devops and 3+ years in DBX. - Strong hands-on experience in Databricks (Spark, Delta Lake, PySpark, MLflow). - Proficiency in SQL and programming languages like Python or Scala. - Experience with Azure cloud data services. - Solid understanding of data modeling, ...
... exceptional candidates from other institutions with demonstrable production AI/ML experience will be considered. Work Experience: 2–4 years of total experience in data/AI engineering, with a minimum of 1 years of hands-on MLOps or LLMOps experience in a production environment. Demonstrated experience deploying and monitoring ML ...
... manage schema evolution, and maintain high data availability . Implement data governance , version control, and CI/CD best practices. Monitor and troubleshoot data pipelines for continuous reliability and efficiency improvements . Why Join KANINI? Join KANINI’s award-winning Data Engineering Team, recognized as the "Outstanding ...
... Experience with NLP techniques and working with LLMs (e.g., Familiarity with prompt engineering, fine-tuning, and model deployment on Azure. - Experience with vector databases (e.g., Azure AI Search, FAISS). MLflow, Azure DevOps). Experience with data visualization tools (e.g., Power BI). Certifications in Azure AI or Data Science. ...
... years of experience in AI/ML, including model development, data preprocessing, EDA, training, and evaluation. - 2+ years of hands‑on experience in Generative AI (LLMs, embeddings, RAG, LLM‑based apps). - 6+ months of hands‑on experience with Agentic AI frameworks (CrewAI / AutoGen / LangGraph / LangChain Agents). - Strong ...
... and responsible AI practices Collaborate with product, data, and platform teams to deliver business solutions Bachelor's or master's degree in computer science, AI, Data Science, or related field Experience 3–5 years of experience in software/ML engineering At least 1–2 years of hands-on experience in Generative AI / LLM-based ...
Role –Gen AI Engineer Location: PAN India Exp: 5+ years Mode Of Interview - F2F Job Description Collect and prepare data for training and evaluating multimodal foundation models. This may involve cleaning and processing text data or creating synthetic data. Develop and optimize large-scale language models like GANs (Generative ...
... Glue, SageMaker), Azure, or GCP. - Strong background in data engineering: ETL/ELT pipelines, data modeling, SQL, and data warehouse technologies (Snowflake, Databricks preferred). - Experience with REST APIs, GraphQL, and enterprise system integrations (CRM, ERP, data marketplaces). - Familiarity with containerization ...
... probability, data mining and data cleaning . - Strong experience with SQL and Python ML packages such as Scikit-learn, XGBoost and Keras. - Experience in building data pipelines and working with at least one cloud platform . - Good knowledge of Git, GitLab and CI/CD . - Familiarity with AI ethics and Responsible AI practices ...
... practical Java experience - 2 years of experience with prompt engineering and prompt/agent scripting - 2 years of hands-on experience integrating third-party AI/ML services and platform APIs (e.g., OpenAI, Anthropic, Salesforce and other SaaS) - Proven ability to design and manage domain-specific languages, prompt templates, ...
... Prompt Engineering, and AI solution design.- Experience designing enterprise-scale AI, data, cloud, and integration architectures.- Hands-on experience with Azure AI Services, Azure OpenAI, AWS AI/ML, GCP AI, or equivalent cloud AI ecosystems.- Strong understanding of MLOps, data pipelines, API integrations, and modern AI application ...
... performance.- Develop and optimize AI-driven extraction workflows using document parsing, chunking, embeddings, RAG, and LLM-based extraction methods.- Deploy and scale AI models on AWS and Azure (SageMaker, Bedrock, Azure AI Foundry) with seamless integration into data pipelines.- Build and maintain CI/CD pipelines for AI model ...
... performance.- Develop and optimize AI-driven extraction workflows using document parsing, chunking, embeddings, RAG, and LLM-based extraction methods.- Deploy and scale AI models on AWS and Azure (SageMaker, Bedrock, Azure AI Foundry) with seamless integration into data pipelines.- Build and maintain CI/CD pipelines for AI model ...