Guide for implementing new data formats in IterableData. Use when adding support for new file formats, compression codecs, or extending format capabilities.
iterable/datatypes/<format>.pyBaseIterable in iterable/base.pyread(), write(), read_bulk(), write_bulk(), etc.iterable/helpers/detect.pytests/test_<format>.pypyproject.toml if neededfrom iterable.base import BaseIterable
class NewFormatIterable(BaseIterable):
def __init__(self, source, mode='r', **kwargs):
super().__init__(source, mode, **kwargs)
# Initialize format-specific resources
def read(self):
# Return iterator of dict objects
pass
def write(self, data):
# Write dict objects to file
pass
def read_bulk(self, size=1000):
# Bulk read for performance
pass
def write_bulk(self, data):
# Bulk write for performance
pass
Update iterable/helpers/detect.py:
detect_file_type() functionExample:
def detect_file_type(filename, content=None):
# Check extension
if filename.endswith('.newformat'):
return 'newformat'
# Check magic numbers
if content and content.startswith(b'MAGIC'):
return 'newformat'
# ... existing detection logic
iterable/codecs/<codec>codec.pyread(), write(), close() methodsiterable/helpers/detect.pypyproject.tomlclass NewCodec:
def __init__(self, fileobj, mode='r'):
self.fileobj = fileobj
self.mode = mode
# Initialize compression library
def read(self, size=-1):
# Decompress and return data
pass
def write(self, data):
# Compress and write data
pass
def close(self):
# Clean up resources
pass
import pytest
from iterable import open_iterable
class TestNewFormat:
def test_read(self):
# Test basic reading
pass
def test_write(self):
# Test basic writing
pass
def test_read_bulk(self):
# Test bulk operations
pass
def test_compressed(self):
# Test with compression (.gz, .bz2, .zst, etc.)
pass
def test_edge_cases(self):
# Empty files, malformed data, etc.
pass
Add to pyproject.toml:
[project.optional-dependencies]
newformat = ["newformat-library>=1.0.0"]
Handle missing dependencies gracefully:
try:
import newformat_library
except ImportError:
raise ImportError(
"newformat support requires 'newformat-library'. "
"Install with: pip install iterabledata[newformat]"
)
Implement capability reporting:
def get_capabilities(self):
return {
'read': True,
'write': True,
'bulk': True,
'totals': False, # Can't count rows without reading
'streaming': True,
'tables': False, # Single table format
}
Look at existing implementations:
iterable/datatypes/csv.py - Text format exampleiterable/datatypes/parquet.py - Binary format exampleiterable/codecs/gzipcodec.py - Compression codec examplewith statementsreset() (or __init__) and raise a clear error (e.g. ReadError) when filename is None.read_bulk() MUST return an empty list []; do not raise StopIteration.