Data Engineer — ML Training Data Pipeline in Hyderabad,… - Jobeax
Vacancy description
Data Engineer — ML Training Data Pipeline in Hyderabad, India
DATAECONOMY
HybridMix of office and remote
₹5 INR
India, Hyderabad
Data Engineer — ML Training Data Pipeline in Hyderabad, India is listed on Jobeax. Browse 30,000+ vacancies available.
Job Title: Data Engineer - ML Training Data Pipeline
Notice period: 0-30 Days
Experience : 5+ Years
Location: Hyderabad OR Pune
We are looking for Data Engineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms raw production traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/test splitting at scale on AWS.
What We Expect:
Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets
Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
Preferred (Not Required): LLM training data prep (chat templates, tool-calling schemas); Axolotl or similar dataset formats; data versioning (DVC, LakeFS); browser-automation trace data or Playwright.
Benefits
Comprehensive Medical Coverage: Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
Robust Protection Plans: Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
Retirement Benefits:
PF and Gratuity provided as per standard government regulations.
Flexible Work Options:
Enjoy hybrid work arrangements & flexible working hours.
Generous Leave Policy: 21 days of annual leave, in addition to 10 company-declared holidays.
Employee Well-being Spaces:
Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.
... skilled Data Engineer to help scale and enhance our internal data observability and analytics platform. This platform integrates with data annotation tools and ML pipelines to provide visibility, insights, and automation across large-scale data operations. You will design and optimize robust data pipelines, build integrations ...
... optimization and data access strategies for APIs serving real-time compliance dashboards, establishing benchmarks and SLOs Lead data infrastructure work for AI/ML features, including dataset curation, feature engineering, and pipeline design supporting cloud-native AI capabilities Define and enforce data quality standards, ...
... Costco Wholesale. Data Engineer The Data Engineer is responsible for developing data pipelines and/or data integrations of for Costco’s enterprise certified data sets that are used for business critical data consumption use cases (i.e. Reporting, Data Science/Machine Learning, Data APIs, etc.). The Data Engineer will partner ...
... Practical exposure to Machine Learning and Deep Learning model development and deployment. - Experience working with large, complex datasets in cloud-based or big data environments. - Practical understanding of Nixtla library is a plus. - Experience with MLOps , model monitoring, and CI/CD pipelines. Domain Knowledge Strong ...
... using LLMs/AI copilots as part of daily workflow (analysis, documentation, communication, process design), not just as a novelty. - Experience working with AI/ML teams and understanding of how annotated data feeds into model training. - Strong understanding of quality frameworks (QA/QC) for spatial data and annotation accuracy. ...
... rigorous analyses, testing assumptions and interpreting complex results. Strong data-wrangling and cleaning skills, including working with large or imperfect datasets, resolving quality issues and developing reproducible data pipelines. Strong machine learning experience, including model selection, feature engineering, ...
... data science libraries such as scikit-learn, pandas, NumPy, XGBoost, LightGBM, PyTorch, TensorFlow, or equivalent frameworks - Experience building end-to-end ML pipelines across data preparation, feature engineering, model training, evaluation, deployment, and monitoring - Hands-on experience with MLOps practices and platforms, ...
... data science and ML, with at least 2 years leading projects or teams in a pharma consulting or client-facing setting. - A Bachelor's or Master's degree in engineering, statistics, mathematics or a relevant quantitative field. - Experience with life sciences data, such as claims (Komodo, IQVIA, Symphony), specialty pharmacy, ...
... Scikit-learn) and statistical analysis. Time series models (Prophet, XGBoost, LSTM for Azure ML). Model performance tracking and retraining strategies using Azure ML. Exposure to AI based solutions in areas of Search, product meta-data generation or discovery. Experience in designing and implementing Data Science best practices ...
... passion for building data architectures that enable smooth and seamless product experiences Are you an all-around data enthusiast with a knack for ETL We're hiring Data Engineers to help build and optimize the foundational architecture of our product's data. We've built a strong data engineering team to date, but have a lot of ...
Big Data Processing: Design and manage scalable data pipelines to process massive datasets efficiently for model training and inference. Build, train, and fine-tune complex neural networks across text, audio, and visual modalities. Cloud Deployment: Architect and deploy models to cloud environments, leveraging distributed ...
... large, multi-source enterprise datasets. - Strong programming proficiency in Python, SQL, and data-modeling techniques for event data. - Experience in process data model design, development, and deployment, including creation of reusable data schemas and pipelines for process analytics. - Experience integrating ML models ...
... production-ready code and perform data analysis with Python and SQL Create workflows for model development and apply feature engineering methods Use Azure AI Search to make data and models easier to consume for business needs Coordinate with developers and project managers using GitLab and Jira Refine data pipelines and tune model performance ...
... is a plus). - Experience with cloud platforms (AWS/GCP/Azure), containerization (Docker/Kubernetes), and geospatial databases (PostGIS) - Strong foundations in ML for agriculture: supervised classification, segmentation, domain adaptation, or transfer learning. - Proficiency in Git, CI/CD pipelines, and frontend testing ...
... workflows, and structured output pipelines for LLM applications. Fine-tune, evaluate, and optimize LLM-powered applications for accuracy, latency, and cost. Implement data preprocessing, feature engineering, and ML model training workflows. Work with structured and unstructured datasets to solve business problems. Collaborate with ...
... hands-on professional experience Education: BE/B.Tech/ME/M.Tech/MCA/MSc IT Primary Skills (must have): Strong Azure or AWS knowledge Strong working knowledge on Data Science, Artificial Intelligence, Machine Learning Secondary Skills (good to have): SQL, Python Professional Attributes: Strong analytical and problem-solving ...
... This includes building secure, observable and human-governed AI-assisted workflows, evaluating their quality, and integrating them responsibly into existing data pipelines and business processes. Data Science, Engineering & Quality Design, build and maintain scalable data-processing and analytical workflows using Python, ...
... opportunities, and business insights. Create dashboards, visualizations, and reports that communicate insights effectively to technical and non-technical stakeholders. Data Engineering & Solution Delivery Collaborate with data engineering teams to establish scalable data pipelines and AI-ready data platforms. Ensure data quality, ...
... Understand data landscape Perform ad-hoc analysis and present results in a clear manner Work on the full lifecycle of machine learning development including sourcing, dataset curation, feature engineering, model training, model tuning, and offline & online experimentation Strong programming skills with minimum 3-6 years of experience ...