Tags
Faster Metrics with Data Marts (overcoming Data Warehouses challenges)
We explain what data marts are and how they help business leaders in managing their company performance by having faster and better metrics. We start by giving a short history of data warehousing and the typical challenges, and then move over to explain data marts and how they can help companies get better metrics today.
What As a Data Mart? And What Challenges Does It Solve?
Data marts simplify access to meaningful business metrics, helping leaders drive performance improvements with clarity and precision. This article breaks down what a data mart is, how it functions, and what makes a metric truly effective. Explore how data marts consolidate information from multiple sources to provide actionable insights, avoid common data pitfalls, and enable smarter decision-making.
A Simple Approach to Master Data Management to Unify Metrics and Insights
Discover the role of master data management (MDM) in achieving consistent and accurate business metrics. This article explains the concept of master data, outlines key challenges organizations face, and introduces two accessible approaches to MDM. By focusing on practical steps and avoiding common pitfalls, we show how businesses can enhance data quality without large budgets or complex systems.
Deploying Prefect on any Cloud Using a Single Virtual Machine
A pattern to improve observability, monitoring and, ultimately, data operations with Prefect. We show how to find the right trade off between number of deployments and improved operations.
dlt and Prefect, a Great Combo for Streamlined Data Ingestion Pipelines
Streamline your data ingestion pipelines with dlt and Prefect. This article explores how combining these powerful open-source tools enables scalable, efficient, and production-ready data workflows. Learn best practices, key features, and real-world insights to simplify data engineering.
Breaking Down Prefect Deployments To Improve The Data Ops Efficiency
Discover how breaking down monolithic ETL flows into modular deployments enhances observability, streamlines troubleshooting, and boosts scalability. Learn to design data pipelines that evolve with your needs while maintaining performance and reliability.
What is a Modular Data Platform?
Learn why modularity is crucial for building scalable, efficient data architectures. This article covers the core components of modern data platforms, from ingestion to governance, and shares best practices for flexibility, interoperability, and security.
Organizing Networking for Data Platforms: Key Connectivity Options
Optimize your data platform by making informed networking decisions. This article explores how networking impacts ELT workflows, covering key connectivity options, security considerations, and best practices. Learn how to design a secure, scalable, and high-performing data platform architecture with the right networking.
How to Setup Data Platform Infrastructure on Google Cloud Platform with Terraform
Learn how to set up a secure, scalable data platform infrastructure on Google Cloud Platform (GCP) using Terraform. This step-by-step guide covers VPC configuration, Compute Engine setup, firewall rules, Identity-Aware Proxy (IAP), Cloud NAT, and more, ensuring a cost-effective, flexible, and secure foundation for your data platform.
Roles in the Context of the Analytics Workflow
Code-based analytics workflows offer a unique advantage - the ability to combine robust collaboration with strict governance. While technically straightforward to implement, this approach only reaches its full potential when aligned with well-adapted business and analytics processes.
Driving Sustainability with Data: Improving CO₂ Emissions Reporting Across Supply Chains
Accurate CO₂ emissions reporting is vital for meeting sustainability goals and regulatory requirements. This article delves into the challenges of Scope 3 emissions, the importance of clean data, and how structured data systems like sustainability marts can improve reporting, ensure compliance, and support better decision-making for businesses.
Flat Tables vs. Snowflake Semantic Models: The Ultimate BI Data Debate
Structuring data for BI is a key decision that impacts performance, scalability, and data consistency. This article compares flat tables and semantic models, highlighting the strengths and trade-offs of each. Learn how a hybrid approach can offer the best of both worlds—combining consistency, flexibility, and efficient analytics across tools and teams.
Getting to Your First Flow Run: Prefect Worker & Deployment Setup
Run your first data ingestion workflow with Prefect, Docker, and Kubernetes. This guide walks through containerized flow execution, Prefect worker deployment, and clean deployment configs, laying the foundation for a scalable, maintainable orchestration layer.
Scaling Secure Data Access: A Systematic RBAC Approach Using Entra ID
Establish scalable, secure access controls for your data platform with a systematic RBAC strategy built on Microsoft Entra ID. This article outlines a five-phase implementation—from user persona mapping to automated auditing—designed to balance flexibility, compliance, and operational efficiency.
CI/CD for Data Workflows: Automating Prefect Deployments with GitHub Actions
The final part of the Data Platform Infrastructure on GCP series covers CI/CD for Prefect deployments using GitHub Actions and Docker. Automate flow builds, worker updates, and streamline orchestration across environments.
SAP Data Ingestion with Python: A Technical Breakdown of Using the SAP RFC Protocol
Streamline SAP data integration with Python by leveraging the RFC protocol. This interview with the lead engineer of a new SAP RFC Connector explores the challenges of large-scale data extraction and explains how a C++ integration improves stability, speed, and reliability for modern data workflows.
Data Platform Cost Optimization: Practical Strategies for Query Performance, Storage, and Cloud Resource Management
Explore how you can dramatically reduce data platform costs without sacrificing performance. This guide breaks down actionable techniques across query tuning, incremental data loading, cloud resource management, and storage lifecycle design.
Why Data Teams Struggle Without Separate Dev and Prod Environments
When development and production share the same data environment, even small changes can trigger costly outages. This article explains why separating dev and prod is foundational for reliable analytics, and how teams can do it without overengineering or blowing the budget.
The Rescue dbt_rerun Deployment: Rebuilding Changed and Broken Models Without Disrupting Production
Keeping production data correct after a dbt change is harder than it looks. Learn how we introduced a dedicated rescue deployment to rebuild exactly what’s needed and when it’s needed, bringing consistency back to production data without costly full reruns or pipeline disruptions.
Running dbt Rescue Rebuild in Production: Operational Playbooks, Failure Models, and Recovery Patterns
Go beyond the setup and into real-world execution. Learn how we run dbt rescue rebuilds in production: scoping dependencies, managing warehouse contention, handling incremental models, and recovering from outages with precision, without introducing new risks to pipeline stability.
What the Claude Code Leak Reveals About Enterprise AI
The Claude Code leak provides a practical look at AI architecture and where value is created in enterprise AI applications. This article explores why governance, workflows, and deterministic logic often matter more than AI itself when building reliable, cost-effective solutions.
Distributing Facts and Dimensions: Governance, Access & Ownership
Building facts and dimensions is only part of the challenge. This article explores how certified data should be distributed across the organization through controlled access paths, ownership models, governance processes, and support structures.
Certified Metrics: From Fact to Dashboard
Certified metrics require more than documented formulas. Learn how facts, measures, dimensions, aggregate metrics, and dashboards work together to create trusted and reusable business metrics.
Hub & Domains: A Practical Data Operating Model
Domains bring data ownership closer to the business, but governance, access management, and shared standards remain difficult to decentralize. This article explores a Hub & Domains operating model that combines business ownership with centralized governance and platform controls.
AI Access Management: Three Governance Layers
AI introduces a new access management challenge: what the user can see, what the agent can see, and what the LLM can see are not the same thing. This article explores how AI governance can be integrated into an existing hub-and-domain data architecture through identity groups, secured schemas, and governed AI harnesses.