Build Real Data Pipelines. Pass Real Interviews. Second Edition — Substantially Expanded. Data engineering is now one of the most in-demand technical careers — but most books either drown beginners in theory or assume cloud experience you don't have yet. This book takes a different path: one running enterprise scenario, built layer by layer, from your first SQL query to a fully orchestrated, production-shaped platform. Written by a practitioner with over fifteen years of enterprise experience, every chapter explains why before how , defines every term the first time it appears, and ends with interview questions, key takeaways, and hands-on exercises. What you will build and master: SQL that interviews test: joins, window functions, CTEs, execution order, and performance tuning with real explain plans - Python for pipelines: pandas, validation, performance optimisation, and two complete portfolio projects - Apache Spark and PySpark: distributed computing concepts, Spark SQL, the DataFrame API, and the Catalyst optimizer - The lakehouse, properly: Parquet, partitioning, and Delta Lake with ACID transactions and time travel - NEW — Dimensional data modeling: star schemas, grain, and slowly changing dimensions — the most heavily tested interview topic - End-to-end projects: a complete batch pipeline and a real-time streaming pipeline, deployed with Docker and Databricks - NEW — Apache Airflow orchestration: DAGs, retries, and backfills that turn individual jobs into one reliable platform - AI-assisted engineering: a deliberate, verification-first practice for using AI coding tools without shipping confidently wrong code What's new in the Second Edition: Two entirely new chapters: Data Modeling and Apache Airflow - Full-colour interior with tips, warnings, and senior-advice callouts throughout - Delta Lake and lakehouse coverage across the book - 70+ interview questions with model answers, a 40-question rapid-fire revision appendix, and a 30-day study plan - Updated certification guidance reflecting current exams from Google Cloud, AWS, Microsoft, and Databricks Who this book is for: Students and graduates targeting a first data engineering role. Testers, analysts, DBAs, and developers transitioning into data. Junior data engineers who want the reasoning behind the tools, not just the syntax. No prior cloud or big data experience required. By the final chapter, you won't just understand data engineering — you'll have built, deployed, and orchestrated it.
| Gtin | 09798284867785 |
| Age_group | ADULT |
| Condition | NEW |
| Gender | UNISEX |
| Product_category | Gl_book |
| Google_product_category | Media > Books |
| Product_type | Books > Subjects > Computers & Technology > Databases & Big Data > Data Warehousing |