Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data in data lakes without ETL pipelines, using embedded DuckDB technology. This capability enables applications to combine operational data with data lake records through familiar PostgreSQL syntax and tools, reducing infrastructure complexity and supporting use cases like real-time dashboards and AI agents.
SpeakesQuery is a local-first search tool for Parquet and SQLite data with optional LLM integration, running entirely on your machine with no cloud dependency or telemetry. It uses a unified query language (SPQL) with semantic search and LLM pipes, installable via Docker with sample data for immediate use.
Aurora PostgreSQL now enables direct querying of Apache Iceberg and Parquet data stored in data lakes without ETL pipelines, using DuckDB's query engine embedded in PostgreSQL. Users can create foreign tables referencing data in Amazon S3 or AWS Glue Data Catalog, allowing existing PostgreSQL applications to access data lake information while optionally materializing it into native Aurora tables.
Amazon Aurora PostgreSQL now supports direct querying of Apache Iceberg and Parquet data stored in data lakes, eliminating the need for ETL pipelines. DuckDB has been embedded into Aurora to enable single queries combining live operational data with historical data lake records using familiar PostgreSQL syntax. The feature is available on Aurora PostgreSQL 17.11+ and 18.6+.
Parquet files offer a simpler, more performant alternative to custom APIs for serving bulk open data. Hosted statically with CORS enabled and queryable via SQL through tools like DuckDB, they eliminate the need for users to learn complex API documentation while enabling efficient data access from browsers, Python, R, and Excel. The author advocates for organizations to serve canonical parquet files at predictable URLs alongside derived formats like CSVs and reports.
duckdb.mk is a minimal Makefile-based build system for local DuckDB projects that automatically parses SQL queries into dependency graphs and generates Parquet tables, requiring only DuckDB, GNU Make, and curl with no configuration needed.
DataHaskell/Dataframe implemented a Parquet writer for Haskell that enables efficient serialization and interoperability with the data science ecosystem. The writeParquet function provides simple usage with defaults, while writeParquetWithOptions allows fine-grained control over row group and page sizes.