You write pipelines. Now own the platform.
AI Data Engineering, GuildTrek’s senior program for engineers who build the data platforms AI runs on. 60 hands-on hours: lakehouse, orchestration, streaming, and the AI/ML data plane, one platform grown end to end. Join the early-access list.
Cohort forming · limited seats
Fees shared on enquiry · nothing due to apply
Instructor-led by Sankar Vema. Know your trainer See the full curriculum, topic by topic What you walk away with
60
Live hours
4
Modules, M0 to M3
3
Practice domains
12+
Tools, one reference each
1
End-to-end data platform
100%
Hands-on, not high-level
Not sure it's for you? Don't take our word for it.
Why is data engineering the real bottleneck for enterprise AI?
Sound familiar?
If any of these is you, keep reading.
If any of these sound like you, this program was built for your next step.
“You ship pipelines, but you have never owned a whole platform.”
You build one end to end, from ingestion to the AI data plane, and defend it.
“You know one tool, but the stack around it keeps shifting.”
Concept first, one reference tool, alternatives named: the skill outlives the tool.
“Your batch jobs work; streaming and real-time still feel like a mystery.”
You build Kafka, CDC, and stream processing with exactly-once, hands-on.
“Everyone wants RAG and features, but the data side is undefined.”
You build embeddings pipelines, vector stores, feature stores, and RAG data.
“Your pipelines pass today and break silently next week.”
Every topic carries idempotency, tests, contracts, lineage, and cost.
“AI writes your pipeline code, but you cannot always trust it.”
You learn to prompt it well and to review what it writes, every topic.
Why this is different
There are plenty of data courses. Few build the AI data plane.
The difference is depth and delivery: one real platform, built the enterprise way, from batch to the data plane AI runs on.
Another data course
GuildTrek
A single-vendor tutorial
Concept, a reference tool, then the field
Disconnected notebook exercises
One platform you grow end to end
Batch only
Batch, streaming, and the AI/ML data plane
Happy-path demos
Idempotency, tests, contracts, lineage, and cost
A PDF certificate
A defended platform plus two more portfolio threads
In 60 hours · part-time, you go from
- “I write pipelines”
- “I know one stack”
- “I have done some batch ETL”
“I designed, built, tested, and defended an end-to-end AI-ready data platform: batch, streaming, and the AI/ML data plane.”
Proof, not a PDF
What you build and defend.
Across the program you grow one platform layer by layer, ending each module with a defended capstone, from a tooled starter to an AI-ready data platform.
The commitment
60 hours · part-time. Live with a trainer, every day.
Live online on weekends: 3 hours on Saturday and 3 hours on Sunday, across ten weekends. Hands-on labs and the offline practice threads happen between sessions.
Saturday
3 hrs
Live session: concept plus a live-lab build on the platform
Sunday
3 hrs
Live session: deeper build, module review, and Q&A
Between sessions
~4 hrs
Your lab work and the IoT / finance practice threads
60 live hours over ten weekends, module by module, each ending in a defended capstone on the commerce platform.
The path
The path, level by level.
Each level ends in a defended capstone. The full topic-by-topic detail is below.
Foundations & toolchain
The batch lakehouse
Streaming & scale
The AI/ML data plane
Total
60 hours
The curriculum
What you'll cover.
A structured, level-by-level path. The full topic-by-topic detail comes on enrolment.
1Module 0 · Foundations & toolchain
~10 hrs- Python and advanced SQL patterns for data engineers
- The toolchain: Git workflow, Docker, cloud, IaC first principles
- Distributed-systems fundamentals: partitioning, replication, consistency
- Storage and file formats: rows vs columns, Parquet, object storage
- The enterprise-quality lens: idempotency, reproducibility, testing, cost
2Module 1 · The batch lakehouse
~20 hrs- Ingestion patterns: batch, API, database sources, incremental loads
- The lakehouse and table formats: Delta, Iceberg (ACID, time travel, schema evolution)
- Distributed processing with Spark and PySpark: execution model, partitions, shuffles, tuning
- Transformation and modeling with dbt; the medallion architecture
- Orchestration with Airflow (reference), Dagster and Prefect (the field)
- Data quality and contracts: tests, expectations, schema enforcement
3Module 2 · Streaming & scale
~15 hrs- Event streaming with Kafka: topics, partitions, delivery semantics
- Change data capture: streaming the operational database in
- Stream processing: Structured Streaming and Flink (windows, watermarks, exactly-once)
- Real-time serving; Kappa vs Lambda; observability and lineage
- Governance, security, and cost at scale: catalogs, access control, PII
4Module 3 · The AI/ML data plane
~15 hrs- Unstructured pipelines: parsing and chunking documents
- Embeddings pipelines and vector databases (pgvector, Qdrant, Milvus)
- RAG data infrastructure: indexing, freshness, retrieval-quality evaluation
- Feature engineering platforms and feature stores; point-in-time correctness
- AI-assisted data engineering: building and reviewing AI-written pipelines
5Every module
- Ships a defended capstone on the commerce platform, plus two offline practice threads (IoT telemetry and financial transactions) that grow the same layer in a different domain
Where it leads
What you'll be ready for.
Cohort forming · limited seats
Stop writing one-off pipelines. Start owning the platform.
Join the early-access list for the first AI Data Engineering cohort.