Data Berlin #48
The calm before the surge
Berlin in August: the U-Bahn has seats, the city is emptying out, and the job board is quieter than usual. But this is the calm before the fall hiring surge, the roles that are open now are the ones companies actually need filled. If you're looking, now's the time to position yourself before everyone comes back from the beach.
If you're staying put, we've picked a few pieces worth your time: from data platform architecture to career strategy, with a side of agentic systems and open-source releases.
📖Interesting Readings
A data-driven reality check on AI energy consumption: data centers use ~1.5% of global electricity, a single ChatGPT query consumes ~0.3 watt-hours, and the real concern is geographic concentration rather than the global total. – Data centers consume around 1.5% of global electricity
A walkthrough of a fully portable analytics stack using DuckDB and DuckLake on Cloudflare R2, with SQLMesh for transformations and GitHub Actions for CI/CD. – A Portable Analytics Stack
Four years after Fundamentals of Data Engineering, Joe Reis asks himself what he'd change. FinOps as a seventh undercurrent? Does AI/ML deserve its own seat at the table? If you've ever used the lifecycle model in your work, this is the author himself stress-testing it in public. – The Data Engineering Lifecycle and Undercurrents, 4 Years Later
"Zero-copy" is one of the most abused terms in data architecture. Alex Merced breaks it into six distinct patterns and honestly traces where the bytes go in each one. Spoiler: three of them make a copy. Read this before your next architecture decision. – What Zero-Copy Actually Costs
Same Hetzner server, same workload, two embedded databases. DuckDB writes 4-15x faster and its read cliff sits 100x further out. If you self-host anything on a small machine, this benchmark gives you actual numbers to decide when to switch. – SQLite vs DuckDB on the same $16 box
A practical decision tree for choosing between full refresh, append, merge, delete+insert, time_interval, and SCD2 strategies based on source behavior rather than destination capabilities. – How to choose an incremental data loading strategy
🚀 Shipping Now
ClickHouse 26.7 — 64 new features, 117 performance optimizations, 300+ bug fixes in the biggest release of the year. Breaking change: S3 access from SQL users no longer resolves the server’s default cloud credentials.
Snowflake 10.25 — Session variable size limit 16KB, object definition size limit 1MB, UUID/DECFLOAT in VARIANT and structured types, Iceberg UNKNOWN data type (Preview). Also: dbt Projects env variables, SERVICE_AGENT user type GA, Multi DT manual refresh.
Databricks July updates — Unit testing for Spark Declarative Pipelines (Beta) with mock data and catalog-table redirection; Lakeflow Designer Summer Release with new Prepare/Enter Data/Select/Visualization operators and improved output (insert/append, materialized views, CSV/Excel/JSON on UC Volumes); Lakebase AI-assisted troubleshooting (Beta) with conversational diagnostics via Genie and telemetry in Unity Catalog.
MotherDuck — Role-based access controls (RBAC) with scoped roles and privileges on Business/Enterprise plans; read/write to Cloudflare R2 as an Iceberg catalog; configurable max runtime for Flights.
OpenAI— GPT Transcribe and GPT Live Transcribe GA for file transcription and streaming; organization and project spend limits for the API; Codex iOS with Mermaid diagrams, interactive forms, and prompt recovery.
Anthropic Claude — Claude Opus 5 launched (near Fable 5 intelligence at half the price); HIPAA self-serve for Enterprise and API; redesigned memory with categorized entries; Claude Cowork on web and mobile; Microsoft 365 connector with write tools.
☕ Upcoming Meetups & Events
Aug 1 — 8x x Bella&Bona Mobile Hack 📱
Aug 3 — Getting European AI Right 🏛️
Aug 6 — Managing Engineers In AI Era 🧠
Aug 6 — Speed VS Sovereignty - Europe's AI Startup Dilemma 🚀
Aug 13 — Python in Berlin BBQ @ c-base 🐍
Aug 27 — Data-Meet-Up in Berlin 📊
Aug TBA — Data Berlin Summer Meetup 🐻
💼 This Week’s Job Picks
📊 Data Analyst
Growth Enablement Internship (w/m/d) – Enpal → Apply here
Senior Data Analyst (f/m/x) – SPARETECH → Apply here
Senior Data Analyst (m/f/d) – FlixMobility → Apply here
SAP Business Intelligence Consultant – Arvato Systems → Apply here
Senior Data Analyst (m/f/d) - Berlin – FREE NOW → Apply here
And more here.
📈 Analytics Engineer
Analytics Engineer (w/m/d) – StackFuel → Apply here
Senior Analytics Engineer - Run & Grow – SumUp → Apply here
Senior Analytics Engineer, Marketing DSA – Vinted → Apply here
Data & Analytics Engineer (m/f/d) – Statista → Apply here
Analytics Engineer (Decision Logic / Rules) – Enpal → Apply here
And more here.
⚙️ Data Engineer
Site Reliability Engineer - Data Platform – N26 → Apply here
Data Engineer (f/m/x) – liveEO → Apply here
Senior Data Engineer, Constellation Services – Planet → Apply here
(Junior) Data Engineer - Data Platform (m/f/d) – 1KOMMA5° → Apply here
Member of Engineering (Data & Analytics) – Poolside → Apply here
And more here.
🤖 AI / ML / LLM + Data Scientist
Principal AI Engineer - Conversational Banking – N26 → Apply here
Staff/Senior AI Engineer, AI for Code – JetBrains → Apply here
Engineer II, ML/Go (Consumer, Global Discovery) – Delivery Hero → Apply here
Agentic AI Specialist (m/w/d) – DocMorris → Apply here
Machine Learning Engineer (f/m/d) – Awin → Apply here
And more here.
That's it for #48. The summer slowdown won't last forever — September always brings a fresh wave of roles, meetups, and conference season. We'll be here when it does.
— Data Berlin team








