Data Engineer — ML Training Data Pipeline in Hyderabad,… - Jobeax
Vacancy description
Data Engineer — ML Training Data Pipeline in Hyderabad, India
DATAECONOMY
HybridMix of office and remote
₹5 INR
India, Hyderabad
Data Engineer — ML Training Data Pipeline in Hyderabad, India is listed on Jobeax. Browse 30,000+ vacancies available.
Job Title: Data Engineer - ML Training Data Pipeline
Notice period: 0-30 Days
Experience : 5+ Years
Location: Hyderabad OR Pune
We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.
What We Expect:
Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets
Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.
Benefits
Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
Retirement Benefits:
PF and Gratuity provided as per standard government regulations.
Flexible Work Options:
Enjoy hybrid work arrangements & flexible working hours.
Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays.
Employee Well-being Spaces:
Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
Position: Data Scientist Experience: 3 – 4 Years We are looking for a Data Scientist to help us see those patterns earlier: which investors need a nudge, which portfolios are drifting off-plan, and which programs are actually moving the needle. Investment and portfolio analytics- build models that surface portfolio drift, ...
... architectures and retrieval-augmented generation (RAG) systems Develop and optimize prompts, evaluation frameworks, and guardrails for LLM-powered applications Engineer scalable data and ML pipelines in Databricks using PySpark, Delta Lake, and MLflow Deploy, monitor, and maintain models in production on Azure (Azure AI Foundry, ...
... experimentation, and prompt design; collaborating with Product and Software Engineering to embed AI/ML into user-facing applications; engaging with DevOps/Platform Engineering on environment setup, CI/CD, monitoring, and reliability; and working with Data Engineering on pipeline design and ingestion strategies. Provide technical ...
... Large Language Models (LLMs). - Expert-level proficiency in Python, SQL and its data science libraries (e.g., Proven track record of leading complex, end-to-end data science projects that have delivered significant business impact. Experience with cloud-based ML platforms / ML ops (e.g., AWS SageMaker, MLflow) and their generative ...
... deploy real-time decisioning rules and ML systems to detect synthetic fraud, digital identity theft, account takeovers, and transactional fraud. - Feature Engineering: Mine complex, large-scale, and alternative data streams (bank transactional data, logs, credit bureau reports, structured/unstructured digital signals) ...
... experimentation, and prompt design; collaborating with Product and Software Engineering to embed AI/ML into user-facing applications; engaging with DevOps/Platform Engineering on environment setup, CI/CD, monitoring, and reliability; and working with Data Engineering on pipeline design and ingestion strategies. Provide technical ...
... scalable solutions Generate actionable insights through data analysis and modeling Deploy, monitor, and improve model performance Qualifications - Bachelor's/Master's in Computer Science, Data Science, Engineering, or related field - 3–8 years of relevant experience Skills: data scientist,genai,python,sql,pyspark,aiml,ml
... our team, choose your own path and work on projects that excite you. The Data and Analytics team drives competitive advantage by providing the best-in-class data, analytics, and insights for the Kmart Group. By using the latest cloud platforms tools, data engineering practices, and data science methods, the team is paving ...
... topic-modeling and trace-clustering approaches (Clio-style summarize-then-embed pipelines, BERTopic, c-TF-IDF) is a strong plus - Proficient in Python and the standard data-science stack (pandas, scikit-learn, statsmodels, numpy) - Data engineering competence — you can design and ship ETL and data models (dbt or equivalent), not ...
... performance.- Collaborate with product and business teams to convert complex data insights into practical business strategies.MUST-HAVE SKILLS :- Data Science, Data Forecasting / Time-Series Forecasting- Python or R, SQL- ARIMA, Exponential Smoothing, Prophet, XGBoost, LightGBM- Pandas, Scikit-learn, Statsmodels, MLflow / ...
... candidate will leverage generative AI and large language models (LLM) to enhance our data-driven decision-making processes. Key responsibilities : Design and develop data models using generative AI and LLM technologies. Implement retrieval-augmented generation techniques to improve data accessibility. Analyze complex datasets to ...
... Mandatory skill — Strong ability to design, evaluate, and improve ML models using robust validation strategies, cross-validation, hyperparameter tuning, feature engineering, and model comparison. Experience selecting appropriate algorithms based on business objectives, data characteristics, interpretability requirements, and ...
... computational engineers, machine learning engineers, software developers, or business representatives across our global organization to research, develop, and deliver data science tools,models, or software for solving challenging business problems in the oil and gas industry. Lead end-to-end delivery of AI/ML solutions: scoping, ...
... cost-effective. help establish and enforce coding standards for data engineering teams. Design, develop and operate high performance, large volume data structures for data-powered products and data science. Implement efficient, distributed and scalable pipelines and integrate data from multiple sources to create data products Implement ...
... Experience with data quality tooling (SODA, Collibra, or similar) Exposure to cloud platforms (Azure, AWS, or GCP) Experience in a regulated or enterprise-scale data environment Prior experience mentoring or leading a small pod of engineers EXPERIENCE - 8+ years in data engineering, with at least 4+ years focused on Snowflake ...
... leading the system design and implementation of technical solutions. Working with data excites you; You have created Big Data architecture, can build and operate data pipelines, and maintain data storage, all within distributed systems. You have a deep understanding of data modeling and experience with modern data engineering ...
... Azure Data Factory, Azure Data Lake, Microsoft Fabric, Logic Apps,Power BI, GitHub, and Data Modeling. Required Skills: - 10+ years of Software Engineering / Data Engineering experience. - Strong SQL Server and T-SQL expertise. - Strong experience working in the Azure environment. - Hands-on experience with Azure Data Factory. ...
... in a modern Snowflake-centred stack, with data quality, testing, and ELT best practices built in from day one. You will also help shape how we use AI within data engineering workflows, from pipeline automation to intelligent data produc Design, build, and scale robust ELT pipelines across ingestion, transformation, modeling, ...
... Engineering, or related field. - 8+ years of experience in Devops and 3+ years in DBX. - Strong hands-on experience in Databricks (Spark, Delta Lake, PySpark, MLflow). - Proficiency in SQL and programming languages like Python or Scala. - Experience with Azure cloud data services. - Solid understanding of data modeling, ...
... on the Bajaj Finance Enterprise Data Platform — a 5PB+ medallion lakehouse built on Azure Databricks and Unity Catalog. This role sits at the intersection of data engineering, model lifecycle management, and AI governance, ensuring that every model — from classical ML to RAG pipelines and autonomous agents — is reproducible, ...