Job Description
Key Responsibilities:
Architect and build end-to-end pipelines using Hadoop, Spark/PySpark, and SQL across batch and real-time systems.
Design and implement data models, storing data efficiently in data lakes, warehouses, or NoSQL databases.
Optimize job performance—tuning Spark, Hive, and SQL workflows; managing resource utilization and partitioning.
Develop, test, and deploy ETL logic, ensuring data quality, lineage, and governance.
Monitor and maintain data platforms, troubleshoot complex production issues, and ensure high availability.
Mentor junior engineers, enforce best practices, and foster knowledge-sharing across engineering teams.
Required Skills & Experience:
5+ years in data engineering; 2+ years specializing in big data frameworks (Hadoop, Spark).
Strong proficiency in PySpark, Hive, SQL, and distributed processing.
Experience with cloud platforms (AWS/GCP/Azure) and orchestration tools (Airflow, Oozie).
Solid understanding of data modeling, partition strategies, and performance tuning.
Familiar with DevOps practices: CI/CD, version control, logging, and monitoring.
Excellent communication skills and ability to collaborate with stakeholders.