Job Title: Data Engineer
Experience: 4+ Years
Location: Noida (Work from Office/Hybrid)
Employment Type: Full-Time
About the Role
We are seeking a talented and experienced Data Engineer with 4+ years of experience in designing, building, and optimizing modern data platforms. The ideal candidate will have expertise in developing scalable data pipelines, ETL/ELT workflows, cloud-based data solutions, and big data technologies. You will work closely with analytics, BI, and machine learning teams to deliver reliable, high-quality data solutions that support business intelligence and advanced analytics.
Key Responsibilities
Design, develop, and maintain scalable and reliable data pipelines for processing large volumes of data.
Build and optimize ETL/ELT workflows to ingest, transform, and load data from multiple sources.
Process and manage both structured and unstructured datasets efficiently.
Develop and maintain robust data models to support reporting, analytics, and business intelligence initiatives.
Optimize data processing performance for speed, scalability, and cost efficiency.
Ensure data quality, integrity, governance, and security across the data platform.
Work with cross-functional teams including Business Intelligence, Analytics, Data Science, and Machine Learning teams to understand data requirements and deliver scalable solutions.
Implement monitoring, troubleshooting, and performance tuning for data pipelines and workflows.
Maintain version control and collaborate using Git-based development practices.
Create technical documentation and follow data engineering best practices.
Required Skills & Experience
4+ years of experience as a Data Engineer or in a similar role.
Strong proficiency in Advanced SQL for querying, optimization, and data transformation.
Hands-on experience with Python for data engineering and automation.
Strong knowledge of Apache Spark for distributed data processing.
Experience working with Databricks for big data analytics and engineering.
Hands-on experience with workflow orchestration tools such as Apache Airflow.
Experience with Apache Kafka for real-time data streaming and event-driven architectures.
Strong understanding of Data Warehousing concepts, dimensional modeling, and data architecture.
Experience with BigQuery or Snowflake as enterprise data warehouse platforms.
Hands-on experience with Google Cloud Platform (GCP) data services, including:
BigQuery
Dataflow
Pub/Sub
Proficiency in Git and version control best practices.
Good to Have
Experience with dbt (Data Build Tool).
Knowledge of the Hadoop ecosystem (HDFS, Hive, etc.).
Experience with Azure Data Factory (ADF).
Exposure to AWS Glue for serverless ETL.
Experience with Terraform for Infrastructure as Code (IaC).
Familiarity with Data Governance and metadata management tools such as Collibra, Alation, Apache Atlas, or similar.