All programs
Coming soon · early-access list open

You write pipelines. Now own the platform.

AI Data Engineering, GuildTrek’s senior program for engineers who build the data platforms AI runs on. 60 hands-on hours: lakehouse, orchestration, streaming, and the AI/ML data plane, one platform grown end to end. Join the early-access list.

Cohort forming · limited seats

Fees shared on enquiry · nothing due to apply

Sankar VemaInstructor-led by Sankar Vema. Know your trainer See the full curriculum, topic by topic

What you walk away with

60

Live hours

4

Modules, M0 to M3

3

Practice domains

12+

Tools, one reference each

1

End-to-end data platform

100%

Hands-on, not high-level

Concept, reference tool, then the field One platform you grow end to end Enterprise + AI-copilot lens on every topic

Not sure it's for you? Don't take our word for it.

Why is data engineering the real bottleneck for enterprise AI?

Explore yourself

Sound familiar?

If any of these is you, keep reading.

If any of these sound like you, this program was built for your next step.

“You ship pipelines, but you have never owned a whole platform.”

You build one end to end, from ingestion to the AI data plane, and defend it.

“You know one tool, but the stack around it keeps shifting.”

Concept first, one reference tool, alternatives named: the skill outlives the tool.

“Your batch jobs work; streaming and real-time still feel like a mystery.”

You build Kafka, CDC, and stream processing with exactly-once, hands-on.

“Everyone wants RAG and features, but the data side is undefined.”

You build embeddings pipelines, vector stores, feature stores, and RAG data.

“Your pipelines pass today and break silently next week.”

Every topic carries idempotency, tests, contracts, lineage, and cost.

“AI writes your pipeline code, but you cannot always trust it.”

You learn to prompt it well and to review what it writes, every topic.

Why this is different

There are plenty of data courses. Few build the AI data plane.

The difference is depth and delivery: one real platform, built the enterprise way, from batch to the data plane AI runs on.

GuildTrek

A single-vendor tutorial

Concept, a reference tool, then the field

Disconnected notebook exercises

One platform you grow end to end

Batch only

Batch, streaming, and the AI/ML data plane

Happy-path demos

Idempotency, tests, contracts, lineage, and cost

A PDF certificate

A defended platform plus two more portfolio threads

In 60 hours · part-time, you go from

  • “I write pipelines”
  • “I know one stack”
  • “I have done some batch ETL”

“I designed, built, tested, and defended an end-to-end AI-ready data platform: batch, streaming, and the AI/ML data plane.”

Proof, not a PDF

What you build and defend.

Across the program you grow one platform layer by layer, ending each module with a defended capstone, from a tooled starter to an AI-ready data platform.

A tooled, tested starter data platform A batch ingestion layer with incremental loads A lakehouse with a table format and time travel A Spark transformation and dbt model layer An orchestrated pipeline with quality gates A Kafka + CDC streaming ingestion A stream-processing job with exactly-once An embeddings + vector index pipeline A feature store with point-in-time correctness A RAG-ready catalog index

The commitment

60 hours · part-time. Live with a trainer, every day.

Live online on weekends: 3 hours on Saturday and 3 hours on Sunday, across ten weekends. Hands-on labs and the offline practice threads happen between sessions.

Saturday

3 hrs

Live session: concept plus a live-lab build on the platform

Sunday

3 hrs

Live session: deeper build, module review, and Q&A

Between sessions

~4 hrs

Your lab work and the IoT / finance practice threads

60 live hours over ten weekends, module by module, each ending in a defended capstone on the commerce platform.

The path

The path, level by level.

Each level ends in a defended capstone. The full topic-by-topic detail is below.

M0

Foundations & toolchain

10hours
M1

The batch lakehouse

20hours
M2

Streaming & scale

15hours
M3

The AI/ML data plane

15hours

Total

60 hours

The curriculum

What you'll cover.

A structured, level-by-level path. The full topic-by-topic detail comes on enrolment.

1Module 0 · Foundations & toolchain

~10 hrs
  • Python and advanced SQL patterns for data engineers
  • The toolchain: Git workflow, Docker, cloud, IaC first principles
  • Distributed-systems fundamentals: partitioning, replication, consistency
  • Storage and file formats: rows vs columns, Parquet, object storage
  • The enterprise-quality lens: idempotency, reproducibility, testing, cost

2Module 1 · The batch lakehouse

~20 hrs
  • Ingestion patterns: batch, API, database sources, incremental loads
  • The lakehouse and table formats: Delta, Iceberg (ACID, time travel, schema evolution)
  • Distributed processing with Spark and PySpark: execution model, partitions, shuffles, tuning
  • Transformation and modeling with dbt; the medallion architecture
  • Orchestration with Airflow (reference), Dagster and Prefect (the field)
  • Data quality and contracts: tests, expectations, schema enforcement

3Module 2 · Streaming & scale

~15 hrs
  • Event streaming with Kafka: topics, partitions, delivery semantics
  • Change data capture: streaming the operational database in
  • Stream processing: Structured Streaming and Flink (windows, watermarks, exactly-once)
  • Real-time serving; Kappa vs Lambda; observability and lineage
  • Governance, security, and cost at scale: catalogs, access control, PII

4Module 3 · The AI/ML data plane

~15 hrs
  • Unstructured pipelines: parsing and chunking documents
  • Embeddings pipelines and vector databases (pgvector, Qdrant, Milvus)
  • RAG data infrastructure: indexing, freshness, retrieval-quality evaluation
  • Feature engineering platforms and feature stores; point-in-time correctness
  • AI-assisted data engineering: building and reviewing AI-written pipelines

5Every module

  • Ships a defended capstone on the commerce platform, plus two offline practice threads (IoT telemetry and financial transactions) that grow the same layer in a different domain

Where it leads

What you'll be ready for.

Data Engineer Analytics Engineer Data Platform Engineer AI / ML Data Engineer

Cohort forming · limited seats

Stop writing one-off pipelines. Start owning the platform.

Join the early-access list for the first AI Data Engineering cohort.