API Guide¶
This guide explains how to choose the right import, backend, and configuration for a task, and how to interpret the results. For exact signatures and docstrings, follow the links to the generated API.
The explicit list of supported public symbols, their canonical imports, and compatibility status is in the Public API Inventory.
Choosing an import¶
fsspeckit re-exports the most common symbols from the root package and from domain packages. Prefer root or domain-package imports; use direct module imports only for backend-specific symbols.
| What you need | Canonical import | Extra |
|---|---|---|
| Create a filesystem | from fsspeckit import filesystem, get_filesystem |
none |
| Filesystem type hints | from fsspeckit import AbstractFileSystem, DirFileSystem |
none |
| AWS / GCP / Azure options | from fsspeckit import AwsStorageOptions, GcsStorageOptions, AzureStorageOptions |
aws, gcp, azure |
| Git provider options | from fsspeckit import GitHubStorageOptions, GitLabStorageOptions |
none |
| Storage-options factory | from fsspeckit.storage_options import from_env, from_dict, merge_storage_options, storage_options_from_uri |
none |
| DuckDB dataset handler | from fsspeckit.datasets.duckdb import DuckDBDatasetIO, create_duckdb_connection |
datasets |
| PyArrow dataset handler | from fsspeckit.datasets.pyarrow import PyarrowDatasetIO |
datasets |
| Schema utilities | from fsspeckit.datasets import cast_schema, convert_large_types_to_normal, opt_dtype_pa, unify_schemas_pa |
datasets |
| Coordinated maintenance | from fsspeckit.core.maintenance import DatasetMaintenanceCoordinator |
datasets |
| Adaptive key tracking | from fsspeckit.datasets.pyarrow import AdaptiveKeyTracker |
datasets |
| SQL filter translation | from fsspeckit.sql import sql2pyarrow_filter, sql2polars_filter |
sql |
| Parallel processing | from fsspeckit.common import run_parallel |
none |
| File synchronization | from fsspeckit.common import sync_dir, sync_files |
none |
| Partition helpers | from fsspeckit.common import get_partitions_from_path, validate_partition_columns |
none |
| Path and credential security | from fsspeckit.common import validate_path, scrub_credentials |
none |
| Logging | from fsspeckit.common import setup_logging, get_logger |
none |
See the Public API Inventory for the full per-symbol breakdown, and Installation for the extras matrix.
Filesystem and core¶
Use filesystem() when you need protocol inference, storage-options objects,
caching, or DirFileSystem path confinement. Use get_filesystem() for a
thinner factory when you already know the protocol and options.
Extended I/O methods (read_json, read_csv, read_parquet, write_json,
write_csv, write_parquet, read_files, write_file, write_files) are
registered on the filesystem instance by importing fsspeckit.core.ext.
Signatures: fsspeckit.core.
Dataset backends¶
fsspeckit offers two dataset backends that share the same handler interface. See Dataset Handlers for the shared contract.
Choosing a backend¶
| Choose DuckDB when | Choose PyArrow when |
|---|---|
You want SQL-based merge logic and ad-hoc parquet_scan queries |
You need streaming merge with explicit memory controls |
| You work in SQL-heavy workflows | You rely on the PyArrow ecosystem and predicate pushdown |
| Your dataset is very large and benefits from DuckDB's optimizer | You are in a memory-constrained environment |
DuckDB¶
The DuckDB backend requires the datasets extra. Create a connection, then the
handler.
Signatures: fsspeckit.datasets.duckdb.
PyArrow¶
The PyArrow backend requires the datasets extra. The handler takes an optional
filesystem.
Signatures: fsspeckit.datasets.pyarrow.
Result interpretation¶
write_dataset() returns a WriteDatasetResult with per-file metadata
(files, total_rows, mode, backend). merge() returns a MergeResult
with row-level counts (inserted, updated, deleted) and file-level lists
(rewritten_files, inserted_files, preserved_files). Use these fields for
auditing and downstream planning rather than re-reading the dataset.
Storage configuration¶
Storage options are structured objects. Configure them directly, from environment variables, or from a URI. See Storage Options for the full guide.
SQL filters¶
Translate SQL WHERE clauses into PyArrow or Polars filter expressions with a
single function. Requires the sql extra.
Signatures: fsspeckit.sql.filters.
Common utilities¶
Cross-cutting helpers live in fsspeckit.common and require no extra
dependency.
Signatures: fsspeckit.common.
Optional dependencies¶
fsspeckit uses lazy imports. Calling a feature whose extra is not installed
raises an ImportError naming the package. Match the error to a row in the
extras matrix.
Related documentation¶
- Dataset Handlers - shared handler interface and backend differences.
- Storage Options - provider configuration.
- Legacy Imports - deprecation mappings.
- How-to Guides - task-oriented recipes.
- Explanation - concepts and architecture.