Big Data Engineer - Hadoop/Spark/Scala in Bengaluru, India - Jobeax
Vacancy description
Big Data Engineer - Hadoop/Spark/Scala in Bengaluru, India
HybridMix of office and remote
India, Bengaluru
Big Data Engineer - Hadoop/Spark/Scala in Bengaluru, India is listed on Jobeax. Browse 30,000+ vacancies available.
Role Overview : We are seeking a seasoned Big Data Engineer to join our high-performing data platform team. In this role, you will be responsible for architecting, developing, and maintaining robust data pipelines that process massive datasets to fuel our analytical engines. You will work closely with cross-functional teams, including data scientists, product managers, and infrastructure engineers, to translate complex business requirements into scalable technical solutions. By optimizing our Hadoop and Spark ecosystems, you will directly influence the speed and reliability of our data-driven decision-making processes, ensuring that our stakeholders have access to high-quality, actionable insights that drive business https://jobeax.com/link/KgVLdM5cThorPSD1 Responsibilities : - Design and implement scalable data pipelines using Scala and Spark to ingest and transform large-scale datasets, ensuring high availability and performance for downstream analytics.- Optimize complex SparkSQL queries and Hadoop jobs to reduce processing latency, directly improving the efficiency of our data infrastructure.- Collaborate with engineering teams to integrate real-time data streams using Kafka, enabling low-latency data availability for critical business applications.- Maintain rigorous data quality standards by implementing automated testing and monitoring frameworks, ensuring reliability for all internal and external data consumers.- Partner with stakeholders to identify data bottlenecks and implement innovative engineering solutions that support the long-term scalability of our Big Data https://jobeax.com/link/1dkmzToQOwmnlB8O Skillset : - Demonstrated expertise in building distributed data systems using Hadoop, Spark, and Scala, with a deep understanding of performance tuning and memory management.- Proven ability to write complex SQL queries and develop efficient data models that support diverse analytical use cases.- Strong proficiency in Python for scripting and automation, complemented by hands-on experience in managing data flows through Kafka.- Exceptional communication skills with the ability to articulate technical concepts to non-technical stakeholders and work effectively within a collaborative, fast-paced environment.- A Bachelors or Masters degree in Computer Science, Engineering, or a related quantitative field, reflecting a strong foundation in data structures and algorithms.- Ability to thrive in a hybrid work model across our Chennai, Bangalore, Pune, or Mumbai offices, demonstrating high levels of self-motivation and adaptability to evolving project requirements.- 5 - 8 years of professional experience in Big Data Engineering, with a track record of delivering high-impact data solutions in enterprise environments. (ref:hirist.tech)
... expertise in PyTorch (preferred) or TensorFlow for building, customizing, and training deep neural networks from scratch. Big Data Ecosystem: Experience with Apache Spark (PySpark) and the Hadoop Ecosystem (HDFS, Hive, MapReduce) for handling, transforming, and querying large-scale distributed datasets. Cloud Architecture: Experience ...
... workflows Strong analytical problem solving and debugging skills Additional Responsibilities: Experience with cloud platforms AWS Azure or GCP Familiarity with big data technologies Spark Hadoop Exposure to workflow schedulers Airflow Prefect Cron Knowledge of CI CD pipelines and version control Git Understanding of data governance ...
... 9-10 Years of years of applicable software engineering experience - Strong fundamentals with experience in Bigdata technologies, Spark, Pyspark, Scala, Pandas, Databricks, Airflow, SQL, - Must have experience in cloud technologies, preferably Microsoft Azure. Must have experience in performance optimization of Spark workloads. ...
JOB DESCRIPTION Location: Pune or Remote (within India) Experience: 5+ years Responsibilities: Design, develop, and maintain high-performance, scalable, and maintainable Scala applications. Develop and maintain RESTful APIs using Scala frameworks Work with relational databases (e.g., MySQL, PostgreSQL) using SQL. Participate ...
... at least four GCP services among Data Flow, Data Proc, Pub Sub, BigQuery, Cloud Functions, Composer, GCS Proficient hands-on programming experience in Spark/Scala (python/java) Proficient in building production level ETL/ELT data pipelines from data ingestion to consumption Data Engineering knowledge (such as Data Lake, ...
... Apache Spark, Hadoop, Azure Data Factory, AWS Glue, GCP Dataflow, Apache Airflow, Data Lakes & Data Warehouses, Git, Linux/Unix, Cloud Migration & Modernization, Data Quality & https://jobeax.com/link/0wGHTJrlTIBvTNSb Profile:- Strong analytical and problem-solving skills.- Good understanding of data engineering and database ...
... Python, Generative AI, LLM and RAG.- Strong understanding of Machine Learning and GenAI algorithms.- Experience with AI cloud platforms - AWS / Azure / GCP.- PySpark and Scala exposure is a plus.- Familiarity with Deep Learning, Big Data environments such as Hadoop/Hive is an advantage.- https://jobeax.com/link/5Im6Nfh21WlEEm2q ...
... data engineering, and DevOps automation. The ideal candidate will manage and optimize Cloudera environments, build CI/CD pipelines, and support enterprise-scale data processing workloads. Key Responsibilities Administer and support Cloudera CDP/CDH platforms, including HDFS, Hive, Spark, YARN, Hue, and CDE. Develop, deploy, ...
... implementing data pipelines using PySpark or Spark Scala- Hands-on experience with Databricks, including development, optimization, and management of large-scale data processing and analytics workload- Must have an understanding of streaming data pipelines for near real-time analytics- Hands-on experience and good understanding ...
... tools.- Strong data modeling knowledge - Fact/Dimension, Normalization, Medallion https://jobeax.com/link/ZwDzpQ6hJ7LcOcF6 to Have : - Finance, Risk, or Compliance data experience.- Spark / Hive / Hadoop experience.- YAML-based pipeline configuration.- CI/CD and Git workflows.- API integrations.- Experience with large-scale production ...
... or Azure environments Python & scripting for automation and data processing Data streaming technologies (Kafka, Spark, Flink) Big Data platforms (Snowflake, Databricks, Hadoop, Teradata) Exposure to AI/ML and advanced analytics Experience with NoSQL databases Ability to design scalable and maintainable data architectures. ...
... software. Environments vary from on-site Hadoop/Spark Clusters to Amazon AWS deployments (EC2, S3, EMR, VPC, Lambda) and Google Cloud deployments (CloudStorage, Dataproc, VPC). In this role you will frequently interact with clients to configure software deployments which transform how they leverage technology for data intelligence. ...
... and lead our push into streaming and real-time fraud detection. We run on AWS and make extensive use of Scala Spark, dbt, and Terraform. Build and expose the data lake/lakehouse so teams can understand business and product performance and make better decisions. Design and evolve a data platform that enables scientists and ...
... Technology, Engineering, or related field. - 8+ years of experience in Devops and 3+ years in DBX. - Strong hands-on experience in Databricks (Spark, Delta Lake, PySpark, MLflow). - Proficiency in SQL and programming languages like Python or Scala. - Experience with Azure cloud data services. - Solid understanding of data modeling, ...
... TensorFlow, PyTorch, Scikit-learn, XGBoost, etc).- Solid understanding of deep learning, NLP, computer vision, recommender systems or generative AI.- Knowledge of big data technologies (Spark, Hadoop) and cloud platforms (AWS, GCP, Azure).- Familiarity with MLOps tools (MLflow, Kubeflow, Airflow, Docker, Kubernetes).- Experience ...
... platform features. Qualifications: - Experience: Minimum of 12 years of hands-on experience in a technical support or engineering role related to Databricks Data Intelligence platform, cloud data platforms, or big data technologies. - Technical Skills: A deep understanding of Databricks architecture and Apache Spark™, ...
Data Engineer Start Date Starts Immediately CTC (ANNUAL) Competitive salary Competitive salary Apply By Not Provided Posted 1 day ago Fresher Job Be an early applicant About the job Requirements Strong understanding of AWS architecture best practices, scalability, security, and cost optimization strategies. Strong hands-on ...
... Engineer Professional Additional Certifications (Preferred) - Databricks Certified Associate Developer for Apache Spark - Cloud platform certifications (Azure Data Engineer Associate, AWS Certified Data Analytics, or Google Cloud Professional Data Engineer) - Relevant data engineering or big data certifications Soft Skills ...
... of data engineering initiatives. Total Experience: 1-4 years in data engineering or related fields. Solid experience with Databricks, focusing on platform engineering, data engineering, and analytics. Technical Skills: Proficiency in Python, Scala, or Java; strong experience with Databricks APIs, Apache Spark, and Delta ...