AWS Big Data Blog
Category: Learning Levels
Configure domain-level VPC networking in Amazon SageMaker Unified Studio
Configuring VPC networking per project across a SageMaker Unified Studio domain creates inconsistent, hard-to-audit networks. This post shows administrators how to configure domain-level VPC networking once, so every new project automatically inherits consistent, private network isolation, then update existing projects and validate connectivity.
Announcing Spark Connect on Amazon EMR on EC2: Interactive PySpark anywhere
Amazon EMR on EC2 now supports Spark Connect, so you can develop and debug PySpark interactively from Amazon SageMaker Unified Studio Data Notebooks or your own IDE while Spark runs on your cluster. This post shows you how to get started from both a Data Notebook and a local IDE.
Query unstructured data in Amazon SageMaker Catalog using generative AI
In Part 2 of this series, sign in as a data consumer, subscribe to enriched unstructured data assets in Amazon SageMaker Catalog, and query them using natural language through a no-code Amazon Bedrock chat agent and Amazon Bedrock model inference.
Best practices for scaling large consumer groups on Amazon MSK
As consumer groups on Amazon MSK scale to thousands of members, the metadata record Kafka writes during rebalances can exceed the 1 MB limit and stall the group. Learn how to estimate metadata size, raise the topic-level limit safely, plan capacity, and apply complementary strategies for scaling large consumer groups.
How Moeve standardized dbt runs across data lakes with Amazon Athena
Moeve standardized how it runs dbt across multiple data lakes by building a centralized, serverless launcher on Amazon Athena, AWS Step Functions, AWS Fargate, Amazon DynamoDB, and Amazon EventBridge, cutting new-project onboarding from days to about 15 minutes while keeping compute close to the data and orchestration loosely coupled.
Discover and govern Snowflake data using SageMaker Unified Studio
Connect Snowflake to Amazon SageMaker Unified Studio to build a unified data catalog. Query federated Snowflake tables without moving data, publish enriched assets to SageMaker Catalog, and validate data quality with AWS Glue Data Quality, all while keeping data in Snowflake.
Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 1: IAM Identity Center (IDC)-based domains
Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using new authentication modes in the Amazon Athena ODBC driver, with no third-party ODBC-JDBC bridge. Part 1 covers IAM Identity Center (IDC)-based domains with both DSN-based and DSN-less connection methods.
Connect Amazon SageMaker Unified Studio to Microsoft Power BI – Part 2: IAM-based domains
Connect Microsoft Power BI directly to governed data in Amazon SageMaker Unified Studio using the Amazon Athena ODBC driver. Part 2 covers IAM-based domains with SageMakerIam authentication, including AWS IAM Identity Center administrator setup, for both DSN-based and DSN-less connection methods.
How United Airlines uses Amazon Redshift and AWS Glue Data Catalog federation to query Databricks-managed data
Learn how United Airlines uses AWS Glue Data Catalog federation to query Databricks Unity Catalog data directly from Amazon Redshift Serverless without duplicating data, using resource links and AWS Lake Formation for governance.
Scale down Kinesis Data Streams on-demand capacity with ODA warm throughput
Amazon Kinesis Data Streams now supports scaling down ingest capacity for on-demand Advantage streams with warm throughput. Learn how the scale-down works, how to monitor stream behavior with Amazon CloudWatch, and best practices for releasing excess capacity after transient traffic bursts.









