Reference library

Python Code Samples

Easy snippets you can copy, study, and run in the browser editor.

5 matches
AI & LLM integration patterns easy

How to Chunk a Long Document for RAG Retrieval in Python

Split text into overlapping chunks at sentence boundaries using a custom Python function suitable for RAG retrieval pipelines.

rag text-chunking nlp
Python
import re
from pathlib import Path

def chunk_document(text, chunk_size=500, overlap=100):
    """Split text into overlapping chunks suitable for RAG retrieval."""
    # Normalize whitespace
    text = re.sub(r'\s+', ' ', text).strip()
    
    chunks = []
    start = 0
    while start < len(text):
        end = min(s…
15 0 Open
ML engineering pipelines easy

How to Load, Save, and Split JSON Data in Python

Provides helper functions to load, save, and split JSON dictionary data for simple ML pipeline preprocessing.

json data-splitting ml-pipeline
Python
import json
from pathlib import Path


def load_json_data(file_path):
    """Load JSON data from a file, returning an empty dict if missing."""
    path = Path(file_path)
    if path.exists():
        with path.open("r", encoding="utf-8") as f:
            return json.load(f)
    return {}


def save_json_data(data, f…
13 0 Open
ML engineering pipelines easy

How to do feature selection with VarianceThreshold in Python

This code demonstrates how to use scikit-learn's VarianceThreshold to remove low-variance features from a NumPy array, keeping only those that vary enough to be useful for modeling.

feature selection sklearn machine learning
Python
import numpy as np
from sklearn.feature_selection import VarianceThreshold

def main():
    # Mock dataset: 4 samples, 5 features
    X = np.array([
        [0.1, 0.2, 1.0, 1.0, 0.5],
        [0.2, 0.2, 0.0, 1.0, 0.4],
        [0.1, 0.2, 1.0, 1.0, 0.6],
        [0.3, 0.2, 1.0, 0.0, 0.5]
    ])

    # Select features w…
14 0 Open
ML engineering pipelines easy

One Hot Encode Categories in Python

Convert a list of categorical strings into one-hot encoded numeric vectors using pure Python and NumPy.

one-hot encoding categorical numpy
Python
import numpy as np

categories = ["red", "green", "blue", "red", "blue", "green", "red"]

unique = sorted(set(categories))
lookup = {cat: i for i, cat in enumerate(unique)}

one_hot = []
for cat in categories:
    row = [0] * len(unique)
    row[lookup[cat]] = 1
    one_hot.append(row)

print("Categories:", categories…
13 0 Open
ML engineering pipelines easy

StandardScaler mock in Python

A pure-Python StandarScaler class that standardizes features to zero mean and unit variance without sklearn.

scaling preprocessing machine-learning
Python
import math

class StandardScaler:
    def __init__(self):
        self.mean_ = None
        self.std_ = None

    def fit(self, X):
        n = len(X)
        self.mean_ = [sum(col) / n for col in zip(*X)]
        self.std_ = []
        for col in zip(*X):
            variance = sum((x - self.mean_[i]) ** 2 for i, x …
12 0 Open

Browse by section

Each section groups closely related Python snippets.

Guide: free Python code samples library

Copy-ready Python snippets for learners and developers

PythonSkillset code samples are short, focused examples organised by topic and difficulty. Every snippet is server-rendered HTML — readable by search engines and easy to copy. Open any sample, read the notes, copy the code, then press Try in editor to run it in the browser with Pyodide.

How to use this library

  1. Pick a topic section — strings, lists, files, functions, and more
  2. Open a sample, read How it works, and copy the code block
  3. Run it in the IDE, tweak values, then take a related quiz or tutorial lesson

Samples vs tutorials and challenges

Samples are quick reference — one concept per page. For step-by-step teaching, use our Python tutorials. To test yourself, try quizzes or coding challenges. Clean up style with the Python formatter.