Amazon Aurora PostgreSQL Enhances Data Access with Iceberg and Parquet Support
Amazon Aurora PostgreSQL now supports direct querying of Iceberg and Parquet data.
On this page 9 sections
Amazon Aurora PostgreSQL has introduced new capabilities that allow users to directly query Apache Iceberg and Parquet data stored in data lakes. This enhancement simplifies data access, enabling the combination of operational and historical data without the need for complex ETL processes, which is significant for businesses looking to streamline their data workflows.
Key takeaways
- Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data.
- This eliminates the need for ETL processes, reducing operational complexity.
- DuckDB integration enhances query performance and efficiency.
- Real-time analytics and AI applications can access both live and historical data seamlessly.
- Users can create foreign tables to facilitate easy querying of external data sources.
What changed in Amazon Aurora PostgreSQL?
Amazon Aurora PostgreSQL has enhanced its functionality by enabling direct querying of data stored in Apache Iceberg and Parquet formats. This change allows users to access and combine operational data with data from their data lakes using existing PostgreSQL applications and tools. According to Amazon's announcement, this capability aims to eliminate the need for traditional extract, transform, and load (ETL) processes, which often involve duplicating data and increasing operational complexity.
Why is this important?
This enhancement matters because it simplifies application development and reduces the infrastructure costs associated with maintaining separate data pipelines. By allowing users to query data directly from data lakes, businesses can streamline their operations and focus on deriving insights from their data rather than managing it.
Benefits of Direct Querying
The ability to directly query Iceberg and Parquet data offers several benefits:
- Reduces operational complexity: Eliminating the need for ETL processes means that businesses can avoid the challenges of data duplication and synchronization.
- Simplifies application development: Users can leverage familiar PostgreSQL syntax, making it easier to integrate new capabilities into existing applications.
- Enables real-time analytics: By accessing both live and historical data in a single query, organizations can enhance their analytics capabilities and support AI applications more effectively.
How does integration with DuckDB improve performance?
DuckDB is now embedded within Aurora PostgreSQL, significantly enhancing query performance. This integration allows users to execute queries that include both live operational data and data from external sources in a single operation. The use of DuckDB optimizes query processing through techniques such as predicate pushdown and column pruning, which improve efficiency by ensuring that only relevant data is read during queries.
What are the performance optimizations?
Some of the key optimizations include:
- Predicate pushdown: This technique reduces the amount of data read by filtering data at the storage level.
- Column pruning: This optimization ensures that only the necessary columns are retrieved, further enhancing query performance.
- Data caching: Frequently accessed data is cached within the Aurora instance, allowing for faster subsequent queries.
How to get started with the new features?
To utilize the new capabilities in Amazon Aurora PostgreSQL, users need to follow a few steps:
- Create an Aurora PostgreSQL cluster and attach an IAM role with the AuroraAnalytics feature.
- Enable the
aurora_analyticsextension within the cluster. - Create foreign tables that point to Iceberg or Parquet data stored in Amazon S3.
- Use familiar PostgreSQL syntax to query the data.
This process is well documented in the Aurora PostgreSQL documentation, making it accessible for users to implement these new features effectively.
What are the use cases and applications?
The integration of Iceberg and Parquet support in Aurora PostgreSQL opens up various use cases:
- Financial transactions: Businesses can integrate recent and historical transaction data for comprehensive reporting and analysis.
- Analytics systems: The support for IRC-compatible catalogs allows for seamless integration with a wide range of analytics systems.
- AI applications: Developers can create AI agents that dynamically access diverse datasets, enhancing the capabilities of their applications.
Frequently asked questions
What is Apache Iceberg?
Apache Iceberg is an open table format for huge analytic datasets that provides features like schema evolution and partitioning, making it easier to manage large amounts of data.
What are the benefits of using Parquet format?
Parquet is a columnar storage file format that is optimized for use with big data processing frameworks, providing efficient data compression and encoding schemes, which improve performance.
How do I create foreign tables in Aurora PostgreSQL?
Users can create foreign tables by using the CREATE FOREIGN TABLE statement, specifying the location of the Iceberg or Parquet data in Amazon S3.
What are the costs associated with using these new features?
There are no additional charges for using the direct querying capabilities; users only pay for the incremental Aurora compute and Amazon S3 request costs for reading data lake files.
Is this feature available in all regions?
Yes, the direct querying of Apache Iceberg and Parquet data is available in all commercial AWS Regions and AWS GovCloud (US) Regions.
Conclusion
Organizations looking to streamline their data access and analytics capabilities should consider leveraging the new features in Amazon Aurora PostgreSQL. By enabling direct querying of Iceberg and Parquet data, businesses can enhance their operational efficiency and drive better insights from their data without the complexity of traditional ETL processes.
Sources
- Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in your data lake | Amazon Web Services — Amazon Web Services (published 30 Sep 2026)