Data pipelines & processing
ETL-style flows, batch transforms, validation, and moving data between formats.
How to Stream a Large JSONL File Line by Line in Python
Process a large JSON-lines file incrementally using streaming techniques to avoid loading the entire file into memory.
import json
def process_large_file(filepath, chunk_size=8192):
"""
Stream a large JSON-lines file line by line, processing each record
without loading the entire file into memory.
"""
total_count = 0
total_sum = 0
with open(filepath, 'r') as f:
while True:
chunk = …
Implement Exactly-Once Transaction Log in Python
A mock transaction log that deduplicates transaction IDs so each is recorded only once, with a dataclass for records and simple in-memory storage.
from dataclasses import dataclass
from typing import Dict, Optional
@dataclass
class TxnRecord:
txn_id: str
status: str
class ExactlyOnceTxnLog:
def __init__(self) -> None:
self._log: Dict[str, TxnRecord] = {}
self._processed_ids: set = set()
def record(self, txn_id: str, status: s…
Browse by section
Each section groups closely related Python snippets.
Data pipelines & processing — Python code examples
What you will find here
This page collects data pipelines & processing snippets — short, copy-ready Python you can paste into our free online IDE and run without installing anything. Each sample includes a plain-English explanation and the full source code.
Samples vs tutorials and challenges
Samples are quick reference — one concept per page. For step-by-step teaching, use our Python tutorials. To test yourself, try quizzes or coding challenges. Clean up style with the Python formatter.