HOW
TO BUILD A
LAKEHOUSE
A modern lakehouse blends data-lake flexibility with warehouse reliability. Let’s break down the essentials and then deploy one on Kubernetes step by step.
Introduction
Every modern organization is flooded with data: logs, API streams, databases, IoT sensors, third-party feeds. The challenge is no longer collecting data, but rather turning it into reliable, accessible, and actionable knowledge.
This is where the Lakehouse comes in: a unified architecture that blends the flexibility of a Data Lake with the governance and performance of a Data Warehouse. In other words: the lakehouse is where your raw JSON files meet enterprise-grade analytics and APIs.
What is a Lakehouse?
At its core, a lakehouse combines two worlds:
- Data Lakes: cheap, scalable object storage (S3, GCS, MinIO) that holds anything: CSVs, Parquet, images, logs.
- Data Warehouses: structured, ACID tables, fast queries, and semantic governance.
A lakehouse unifies them with a transactional table format (Apache Iceberg, Delta Lake, Apache Hudi) on top of object storage — delivering scalability, reliability, and flexibility in one place.
Data Warehouse
Data Lake
Lakehouse
→ Single governed access point for BI + DS + ML
So to conclude:
Data Lake
Cheap, scalable object storage (S3, GCS, MinIO) for any format (CSV, Parquet, images, logs…)
Data Warehouse
Structured, governed tables with ACID, semantic models, and fast SQL.
Lakehouse
Transactional tables on object storage via (Apache Iceberg, Delta Lake, Apache Hudi) — one platform for BI and ML.
What Makes a Good Lakehouse?
A "good" Lakehouse is not just storage. It's a living ecosystem with several key qualities:
Open & Interoperable
• Based on open formats (Parquet, Iceberg) → no vendor lock-in
• Works with multiple query engines (Trino, Spark, Dask, Flink)
Layered Architecture (Bronze → Silver → Gold)
• Bronze: raw ingested data (no assumptions)
• Silver: cleaned, typed, deduplicated
• Gold: aggregated, business-ready
• This layering enforces trust and avoids "data swamp" syndrome
Automated Ingestion & Orchestration
• No-code tools (Airbyte, Fivetran) for fast onboarding
• Orchestration (Argo Workflows, Airflow) for reproducibility
Unified Query & Serving
• Trino/Presto for federated SQL across Iceberg, Postgres, ClickHouse
• APIs (Hasura, PostgREST, GraphQL) for apps and services
Governance & Security
• Centralized auth (Keycloak, Azure AD, Okta)
• Fine-grained permissions, lineage, metadata (Nessie, Amundsen, DataHub)
Observability & Reliability
• Metrics, logging, monitoring (Prometheus, Grafana, Loki)
• Automated backups and recovery (Velero, S3 versioning)
Data Flow Pipeline
From Raw to Bronze, then Silver, then Gold.
Building Blocks of a Modern Lakehouse
Here's what a minimal but powerful stack can look like (deployed on Kubernetes):
- • Object Storage: MinIO (S3-compatible) for bronze/silver/gold zones.
- • Table Format: Apache Iceberg + Project Nessie (Git-like catalog).
- • Structured Stores: PostgreSQL (metadata, serving), ClickHouse (real-time analytics).
- • Query Engine: Trino for unified SQL across sources.
- • Ingestion: Airbyte (no-code connectors to S3/PG/CH).
- • Orchestration: Argo Workflows for ETL jobs.
- • API Layer: APISIX + Hasura (REST/GraphQL).
- • BI / Viz: Metabase or Superset.
- • Security: Keycloak for SSO, JWT, LDAP.
- • Observability: Prometheus + Grafana + Loki.
Architecture Flow
Interactive architecture diagram showing the complete Lakehouse data flow with modern open-source technologies

Kubernetes
What is Kubernetes?
Kubernetes (K8s) is an open-source container orchestration platform that automates the deployment, scaling, and management of containerized applications. Originally designed by Google and now maintained by the Cloud Native Computing Foundation, Kubernetes provides a robust and scalable environment for running distributed workloads.
Core Features
- • Automated Orchestration: Intelligent container management
- • Auto-scaling: Horizontal and vertical scaling on demand
- • Self-healing: Automatic restart of failed containers
- • Load Balancing: Intelligent traffic distribution
- • Rolling Updates: Zero-downtime deployments and rollbacks
Distributed Architecture
- • Pods: Fundamental deployment units
- • Services: Stable network abstraction
- • Deployments: Declarative application management
- • ConfigMaps & Secrets: Secure configuration management
- • Persistent Volumes: Durable and shared storage
Why Kubernetes for our Lakehouse?
Native Scalability
Lakehouse workloads require dynamic scaling based on processing needs. Kubernetes enables automatic scaling of:
- • Spark executors based on workload
- • Query services (Trino/Presto)
- • Data ingestion pipelines
- • Metadata services
Multi-Tenancy
Secure isolation of environments and teams on shared physical infrastructure:
- • Namespaces for logical isolation
- • RBAC for granular access control
- • Network policies for network security
- • Resource quotas per team
Complex Orchestration
Sophisticated management of dependencies and data workflows:
- • Jobs and CronJobs for ETL
- • Operators for Apache Spark, Kafka
- • Service mesh for communication
- • Built-in monitoring and observability
Lakehouse-Specific Benefits
🚀 Performance & Efficiency
- • Dynamic compute resource allocation
- • Automatic workload optimization
- • Intelligent cache and memory management
- • Native parallel processing
🔒 Security & Governance
- • Encryption in transit and at rest
- • Complete audit trails for access
- • Integration with existing IAM systems
- • Automated compliance (GDPR, SOX, etc.)
🔧 DevOps Operations
- • GitOps for configuration as code
- • Native CI/CD pipelines
- • Rolling updates without downtime
- • Automatic rollback on errors
💰 Cost Optimization
- • Intelligent resource sharing
- • Spot instances for batch workloads
- • Autoscaling to reduce idle costs
- • Detailed consumption metrics
🎯 Key Takeaway
Kubernetes transforms our lakehouse into a resilient, cloud-native platform where each component (ingestion, storage, processing, analytics) can be deployed, updated, and scaled independently. This microservices approach ensures high availability,elastic scalability, and simplified maintenance, while optimizing infrastructure costs.
Deploying Your Lakehouse
Complete Kubernetes Implementation Guide
This hands-on guide shows how to build a modern, interoperable lakehouse on Kubernetes using a pragmatic, minimal stack. We'll deploy a complete data platform with object storage, table formats, query engines, ingestion pipelines, APIs, and observability, enhanced with a comprehensive AI/ML stack for modern data science workflows.
🏗️ Architecture Overview
Storage Layer
- • MinIO (S3-compatible) for bronze/silver/gold zones
- • Apache Iceberg + Project Nessie (Git-like catalog)
- • PostgreSQL & ClickHouse for structured stores
Processing Layer
- • Trino (unified SQL across all sources)
- • Airbyte (no-code connectors)
- • Argo Workflows (ETL/ELT orchestration)
API & Gateway
- • APISIX (ingress + gateway)
- • Hasura (instant GraphQL on Postgres)
- • NocoDB (Airtable-like UI on Postgres)
AI/ML Stack
- • JupyterHub (multi-user notebooks)
- • MLflow (experiments & model registry)
- • vLLM (OpenAI-compatible LLM serving)
🤖 Enhanced AI/ML Capabilities
Our lakehouse includes a complete AI/ML platform that transforms your data into intelligent applications:
🧠 AI Coding & MLOps
- • JupyterHub: Multi-user notebook environment
- • MLflow: Experiment tracking & model registry
- • Native integration with lakehouse data
📊 No-Code Data Apps
- • NocoDB: Airtable-like UI on Postgres
- • Visual data management & CRUD operations
- • Perfect for analysts and business users
🚀 High-Performance LLM
- • vLLM: GPU-optimized LLM serving
- • OpenAI-compatible API endpoints
- • Secured by APISIX with authentication
💡 Complete Integration: Build RAG applications using your lakehouse data, track ML experiments with artifact storage in MinIO, and create interactive dashboards for business users - all within a unified, cloud-native platform.
Prerequisites
Required Infrastructure
- • Kubernetes v1.26+ with default StorageClass
- • kubectl & helm installed and configured
- • Ingress controller (we'll use APISIX)
- • cert-manager (optional, for TLS via ACME)
- • DNS entries pointing to your load balancer
Step 1: Create Namespaces
kubectl create ns data
kubectl create ns ml
kubectl create ns gateway
kubectl create ns security
kubectl create ns observability
kubectl create ns biObject Storage: MinIO
S3-compatible object storage that will serve as our data lake foundation. We'll create bronze/silver/gold buckets plus dedicated buckets for Airbyte, Iceberg metadata, and MLflow artifacts.
Deploy MinIO
helm repo add minio https://charts.min.io/
helm repo update
helm upgrade --install minio minio/minio \
--namespace data \
--set mode=standalone \
--set rootUser=minioadmin,rootPassword=minioadmin \
--set persistence.size=500Gi \
--set resources.requests.memory=1GiBootstrap Buckets
apiVersion: batch/v1
kind: Job
metadata:
name: minio-bootstrap
namespace: data
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: mc
image: minio/mc:latest
env:
- name: MINIO_ROOT_USER
value: minioadmin
- name: MINIO_ROOT_PASSWORD
value: minioadmin
command: ["sh","-c"]
args:
- |
set -e
mc alias set local http://minio.data.svc.cluster.local:9000 \$MINIO_ROOT_USER \$MINIO_ROOT_PASSWORD
for b in bronze silver gold airbyte iceberg mlflow; do
mc mb -p local/\$b || true
done
mc policy set public local/bronze || true📝 Connection Details
- • Endpoint:
http://minio.data.svc.cluster.local:9000 - • Access Key:
minioadmin - • Secret Key:
minioadmin(change in production!) - • Region:
us-east-1
Structured Stores: PostgreSQL & ClickHouse
PostgreSQL
Primary database for metadata, serving layer, and transactional workloads.
helm repo add bitnami https://charts.bitnami.com/bitnami
helm upgrade --install postgres bitnami/postgresql \
-n data \
--set auth.postgresPassword=postgres \
--set primary.persistence.size=50GiClickHouse
High-performance columnar database for real-time analytics and OLAP queries.
helm repo add clickhouse https://charts.clickhouse.com/
helm upgrade --install clickhouse clickhouse/clickhouse \
-n data \
--set shards=1,replicas=1 \
--set storage.size=200GiIceberg Catalog: Nessie
Nessie provides Git-like versioned catalogs for Apache Iceberg, enabling time-travel queries, branching, and metadata versioning.
apiVersion: apps/v1
kind: Deployment
metadata:
name: nessie
namespace: data
spec:
replicas: 1
selector:
matchLabels: { app: nessie }
template:
metadata:
labels: { app: nessie }
spec:
containers:
- name: nessie
image: ghcr.io/projectnessie/nessie:latest
ports:
- containerPort: 19120
env:
- name: QUARKUS_DATASOURCE_JDBC_URL
value: jdbc:postgresql://postgres-postgresql.data.svc.cluster.local:5432/postgres
- name: QUARKUS_DATASOURCE_USERNAME
value: postgres
- name: QUARKUS_DATASOURCE_PASSWORD
value: postgres
---
apiVersion: v1
kind: Service
metadata:
name: nessie
namespace: data
spec:
selector: { app: nessie }
ports:
- port: 19120
targetPort: 19120Service URL: http://nessie.data.svc.cluster.local:19120/api/v2
Query Engine: Trino
Trino acts as our SQL virtual warehouse, providing unified query access across Iceberg, PostgreSQL, ClickHouse, and object storage.
Deploy Trino
helm repo add trino https://trinodb.github.io/charts/
helm upgrade --install trino trino/trino -n data \
--set server.workers=2 \
--set service.type=ClusterIPConfigure Catalogs
apiVersion: v1
kind: ConfigMap
metadata:
name: trino-catalogs
namespace: data
labels: { app: trino }
data:
iceberg.properties: |
connector.name=iceberg
catalog.warehouse=s3://iceberg
iceberg.catalog.type=nessie
iceberg.nessie.uri=http://nessie.data.svc.cluster.local:19120/api/v2
fs.native-s3.enabled=true
s3.endpoint=http://minio.data.svc.cluster.local:9000
s3.path-style-access=true
s3.aws-access-key=minioadmin
s3.aws-secret-key=minioadmin
clickhouse.properties: |
connector.name=clickhouse
clickhouse.http-port=8123
connection-url=jdbc:clickhouse://clickhouse-clickhouse.data.svc.cluster.local:8123
postgres.properties: |
connector.name=postgresql
connection-url=jdbc:postgresql://postgres-postgresql.data.svc.cluster.local:5432/postgres
connection-user=postgres
connection-password=postgresData Ingestion: Airbyte
No-code data integration platform for moving data from various sources into our lakehouse.
helm repo add airbyte https://airbytehq.github.io/helm-charts
helm upgrade --install airbyte airbyte/airbyte -n data \
--set webapp.service.type=ClusterIP🔧 Configuration
After deployment, access Airbyte UI and configure destinations:
- • S3 (MinIO): Use MinIO endpoint, enable path-style access
- • PostgreSQL: Connect to our Postgres service
- • ClickHouse: HTTP endpoint on port 8123
ETL Orchestration: Argo Workflows
Container-native workflow engine for orchestrating data transformation pipelines.
helm repo add argo https://argoproj.github.io/argo-helm
helm upgrade --install argo-wf argo/argo-workflows -n data \
--set server.service.type=ClusterIPExample: Bronze to Silver ETL
This workflow transforms raw CSV data from bronze bucket into a partitioned Iceberg table:
apiVersion: argoproj.io/v1alpha1
kind: Workflow
metadata:
name: bronze-to-silver-iceberg
namespace: data
spec:
entrypoint: etl
templates:
- name: etl
steps:
- - name: bronze-to-silver
template: spark-sql
- name: spark-sql
container:
image: apache/spark:3.5.1
command: ["/bin/sh","-lc"]
args:
- |
spark-sql \
--packages org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.6.1,org.apache.hadoop:hadoop-aws:3.3.4,software.amazon.awssdk:bundle:2.20.156 \
--conf spark.sql.catalog.iceberg=org.apache.iceberg.spark.SparkCatalog \
--conf spark.sql.catalog.iceberg.catalog-impl=org.apache.iceberg.nessie.NessieCatalog \
--conf spark.sql.catalog.iceberg.uri=http://nessie.data.svc.cluster.local:19120/api/v2 \
--conf spark.sql.catalog.iceberg.warehouse=s3a://iceberg \
--conf spark.sql.catalog.iceberg.ref=main \
--conf spark.hadoop.fs.s3a.endpoint=http://minio.data.svc.cluster.local:9000 \
--conf spark.hadoop.fs.s3a.path.style.access=true \
--conf spark.hadoop.fs.s3a.access.key=minioadmin \
--conf spark.hadoop.fs.s3a.secret.key=minioadmin <<'SQL'
CREATE SCHEMA IF NOT EXISTS iceberg.silver;
CREATE TABLE IF NOT EXISTS iceberg.silver.nyc_trips (
vendor_id string,
tpep_pickup_timestamp timestamp,
tpep_dropoff_timestamp timestamp,
passenger_count int,
trip_distance double,
fare_amount double
)
USING iceberg
PARTITIONED BY (days(tpep_pickup_timestamp));
CREATE OR REPLACE TEMPORARY VIEW bronze_csv
USING csv
OPTIONS (
paths 's3a://bronze/nyc_trips/*.csv',
header 'true',
inferSchema 'true'
);
INSERT INTO iceberg.silver.nyc_trips
SELECT vendor_id,
cast(tpep_pickup_timestamp as timestamp) as tpep_pickup_timestamp,
cast(tpep_dropoff_timestamp as timestamp) as tpep_dropoff_timestamp,
cast(passenger_count as int) as passenger_count,
cast(trip_distance as double) as trip_distance,
cast(fare_amount as double) as fare_amount
FROM bronze_csv;
SQLAPI Layer: APISIX + Hasura
APISIX Gateway
High-performance API gateway for ingress and service mesh capabilities.
helm repo add apisix https://charts.apiseven.com
helm upgrade --install apisix apisix/apisix -n gateway \
--set gateway.type=LoadBalancerHasura GraphQL
Instant GraphQL APIs over PostgreSQL with real-time subscriptions.
helm repo add hasura https://hasura.github.io/helm-charts
helm upgrade --install hasura hasura/graphql-engine -n gateway \
--set graphqlEngine.env[0].name=HASURA_GRAPHQL_DATABASE_URL \
--set graphqlEngine.env[0].value=postgres://postgres:[email protected]:5432/postgres \
--set graphqlEngine.env[1].name=HASURA_GRAPHQL_ENABLE_CONSOLE \
--set graphqlEngine.env[1].value=trueBusiness Intelligence: Metabase
Open-source BI platform for creating dashboards and exploring data across all our sources.
helm repo add metabase https://helm.metabase.com
helm upgrade --install metabase metabase/metabase -n bi \
--set database.type=postgres \
--set database.encryption.enabled=false \
--set database.connectionURI=postgres://postgres:[email protected]:5432/postgres📊 Data Source Connections
Connect Metabase to PostgreSQL, ClickHouse, and Trino for comprehensive analytics across your entire lakehouse.
AI & ML Platform
JupyterHub
Multi-user notebook environment for data science and ML development.
helm repo add jupyterhub https://jupyterhub.github.io/helm-chart/
helm upgrade --install hub jupyterhub/jupyterhub -n ml \
--set proxy.service.type=ClusterIP \
--set singleuser.image.name=python \
--set singleuser.image.tag=3.11-slimMLflow
ML lifecycle management with experiment tracking and model registry.
apiVersion: apps/v1
kind: Deployment
metadata:
name: mlflow
namespace: ml
spec:
replicas: 1
selector:
matchLabels: { app: mlflow }
template:
metadata:
labels: { app: mlflow }
spec:
containers:
- name: mlflow
image: ghcr.io/mlflow/mlflow:latest
args: ["mlflow","server","--host","0.0.0.0","--port","5000",
"--backend-store-uri","postgresql+psycopg2://postgres:[email protected]:5432/postgres",
"--default-artifact-root","s3://mlflow/"]
env:
- name: MLFLOW_S3_ENDPOINT_URL
value: http://minio.data.svc.cluster.local:9000
- name: AWS_ACCESS_KEY_ID
value: minioadmin
- name: AWS_SECRET_ACCESS_KEY
value: minioadmin
ports:
- containerPort: 5000vLLM: High-Performance LLM Serving
OpenAI-compatible API for serving large language models with high throughput and GPU acceleration.
apiVersion: apps/v1
kind: Deployment
metadata:
name: vllm
namespace: ml
spec:
replicas: 1
selector:
matchLabels: { app: vllm }
template:
metadata:
labels: { app: vllm }
spec:
containers:
- name: vllm
image: vllm/vllm-openai:latest
args:
- "--model=/models/Qwen2.5-7B-Instruct"
- "--max-num-seqs=64"
- "--dtype=float16"
ports:
- containerPort: 8000
resources:
limits:
nvidia.com/gpu: 1
volumeMounts:
- name: models
mountPath: /models
volumes:
- name: models
persistentVolumeClaim:
claimName: models-pvc
---
apiVersion: v1
kind: Service
metadata:
name: vllm
namespace: ml
spec:
selector: { app: vllm }
ports:
- port: 8000
targetPort: 8000Security & Observability
Keycloak SSO
Identity and access management with OpenID Connect and JWT.
helm repo add codecentric https://codecentric.github.io/helm-charts
helm upgrade --install keycloak codecentric/keycloak -n security \
--set keycloak.username=admin,keycloak.password=adminPrometheus + Grafana
Complete observability stack with metrics, logs, and dashboards.
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo add grafana https://grafana.github.io/helm-charts
helm upgrade --install kube-prom prometheus-community/kube-prometheus-stack -n observability
helm upgrade --install loki grafana/loki-stack -n observability \
--set promtail.enabled=trueEnd-to-End Data Workflow
🔄 Complete Data Pipeline
Ingest Raw Data
Use Airbyte to create connectors from various sources (APIs, databases, files) and load raw data into the bronze bucket in MinIO.
Transform to Silver
Submit Argo Workflows to process bronze data using Trino SQL, creating clean, validated Iceberg tables in the silver layer with proper partitioning.
Aggregate to Gold
Create business-ready aggregated tables and materialized views in PostgreSQL and ClickHouse for fast serving and analytics.
Serve & Visualize
Expose data via Hasura GraphQL APIs, create Metabase dashboards, and enable ML workflows through JupyterHub with MLflow tracking.
AI Integration
Use vLLM for LLM inference, integrate with data for RAG applications, and maintain comprehensive observability with Prometheus and Grafana.
Sample Queries & Operations
Query Iceberg via Trino
-- Connect to Trino
trino --server trino.data.svc.cluster.local:8080 \
--catalog iceberg --schema silver
-- Query partitioned data with time travel
SELECT date(tpep_pickup_timestamp) d, count(*)
FROM nyc_trips
WHERE tpep_pickup_timestamp >= TIMESTAMP '2024-01-01'
GROUP BY 1
ORDER BY 1 DESC
LIMIT 10;
-- Create gold aggregation
CREATE SCHEMA IF NOT EXISTS iceberg.gold;
CREATE TABLE IF NOT EXISTS iceberg.gold.daily_trips AS
SELECT date(tpep_pickup_timestamp) d, count(*) trips
FROM iceberg.silver.nyc_trips
GROUP BY 1;Test vLLM API
curl -s https://llm.example.local/v1/chat/completions \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model":"Qwen2.5-7B-Instruct",
"messages":[{"role":"user","content":"Summarize yesterday'''s lakehouse loads."}],
"temperature":0.2
}'Production Considerations
🔒 Security Hardening
- • Rotate credentials and use Kubernetes Secrets
- • Enable encryption at rest (MinIO KMS, database TDE)
- • Configure network policies and service mesh
- • Implement RBAC and pod security standards
- • Set up certificate management with cert-manager
⚡ Performance & Scale
- • Scale Trino workers based on query load
- • Implement MinIO distributed mode with erasure coding
- • Use ClickHouse clusters with replication
- • Optimize Iceberg table maintenance (OPTIMIZE, VACUUM)
- • Configure resource quotas and limits
📊 Monitoring & Alerts
- • Create Grafana dashboards for all services
- • Set up alerting rules for data quality and SLAs
- • Monitor pipeline health and data freshness
- • Track resource utilization and costs
- • Implement log aggregation and analysis
🔄 Data Operations
- • Implement data quality checks with Great Expectations
- • Set up automated backup and disaster recovery
- • Create data lineage tracking and cataloging
- • Establish data governance and compliance processes
- • Plan for schema evolution and migration strategies
🎉 Congratulations!
You now have a complete, production-ready lakehouse running on Kubernetes. This modern data platform combines the flexibility of data lakes with the performance of data warehouses, enhanced with AI/ML capabilities and comprehensive observability.