AWS Big Data Blog
Category: Analytics
How Sony LIV built real-time video streaming analytics with AWS
Real-time data analytics is transforming how streaming platforms serve their audiences. Learn how Sony LIV built a comprehensive, real-time streaming analytics solution on AWS using Amazon Kinesis Data Streams, Amazon EMR, and Apache Iceberg.
Network connectivity patterns for the next generation of Amazon OpenSearch Serverless
The next generation of Amazon OpenSearch Serverless uses standard AWS PrivateLink endpoints on the on.aws domain. This post shows nine connectivity patterns for private access, from a single VPC to multiple VPCs, cross-account, on-premises, and cross-Region, with the DNS resolution and data path for each.
How Moovit achieved 33% cost optimization through architectural modernization
Learn how Moovit modernized its data platform with a multi-engine lakehouse architecture: offloading heavy aggregation workloads from Amazon Redshift to Amazon EMR with Spark SQL, isolating workloads with Amazon Redshift Serverless, and cutting overall data pipeline cost by 33%.
Query Amazon S3 Tables from Amazon EMR Trino using the Iceberg REST endpoint
Learn how to query Amazon S3 Tables from Trino on Amazon EMR using the Apache Iceberg REST catalog endpoint. This post shows how to deploy the integration with AWS CloudFormation, configure the Trino catalog, and run SQL to create, query, and manage Apache Iceberg tables.
Build a dynamic streaming data lake with Apache Iceberg and Apache Flink
Learn how to build a dynamic streaming data lake on Amazon Managed Service for Apache Flink that adapts to new event types and schema changes without stopping the pipeline, using Apache Iceberg’s Dynamic Iceberg Sink for per-record table routing and automatic schema evolution.
Observing and evaluating production agents using OpenSearch Agent Health
Learn how to observe and evaluate production AI agents by combining an agent running on AWS with OpenSearch Agent Health. This post walks through deploying an agent and its observability pipeline to AWS, then using Agent Health to explore traces and run evaluations that measure and improve agent quality over time.
Accelerate Apache Spark debugging on Amazon EMR with AWS DevOps Agent
Extend AWS DevOps Agent to investigate Apache Spark failures on Amazon EMR. This post shows how to register the Apache Spark Troubleshooting Agent for Amazon EMR as a custom MCP capability provider over AWS PrivateLink, so a single agent chat session diagnoses a failing Spark job from an Amazon CloudWatch alarm to a line-numbered root cause.
Deliver real-time data to streaming tables for Apache Iceberg with Amazon Kinesis Data Streams
Amazon Kinesis Data Streams now supports streaming tables, a fully managed capability that continuously delivers your streaming data as queryable Apache Iceberg tables on Amazon S3 Tables. Streaming tables reduce data delivery costs to S3 Tables by up to 50% compared to self-managed alternatives and reduce downstream query costs by up to 30% through intelligent inline compaction that eliminates the small file problem. You need no custom applications, no self-managed compute, and no operational overhead.
Measuring and improving search quality with Amazon OpenSearch Service
Most teams struggle to answer a deceptively simple question: is my search returning relevant results? This post shows how to capture User Behavior Insights (UBI) data on Amazon OpenSearch Service and use Search Relevance Workbench (SRW) to turn those signals into relevance judgments and evaluate search quality.
Build a real-time event pipeline with Spark Real-Time Mode on AWS Glue 6.0
With AWS Glue 6.0, you can build real-time, near-real-time, and batch data pipelines on a single platform. Using a financial market-risk example, learn how to flag high-risk trades with sub-second latency using Spark Real-Time Mode, store heterogeneous pricing vectors with Apache Iceberg v3 Variant columns, and run batch analytics with Arrow-native UDFs.









