Courses/Python Mastery/Module 10: Power Tools
Module 10 · Lesson 630 minBeginner

🧪 Lab: File Organizer

What you'll build
Apply power tools: comprehensions, pathlib, and the standard library in a useful tool.

Introduction

If your Downloads folder looks like a digital landfill — a chaotic mix of PDFs, screenshots, ZIP archives, Excel files, and random executables — you are not alone. Manually sorting files is one of those tedious tasks that screams out for automation.

In this guide you will build a File Organizer web app using Python and Streamlit. The app scans any folder on your computer, groups files by their type, shows you a preview of the planned moves, and then sorts everything into clean, categorised subfolders with one click. A built-in undo log lets you reverse the operation if needed.

This project complements other automation-focused tutorials on this blog such as the Password Generator and the Weather App — both great examples of turning a Python script into a useful interactive tool.


What You Will Build

The finished app will:

  • Accept a folder path as input
  • Scan and list all files with their sizes and extensions
  • Map each extension to a category (Images, Documents, Videos, etc.)
  • Show a preview table of planned moves before anything changes
  • Move files into category subfolders on confirmation
  • Keep an undo log so you can revert in one click
  • Support a dry-run mode for safe previewing

  • Prerequisites

  • Python 3.8+ — Download from python.org.
  • Streamlit — Install with pip:
  • code
    pip install streamlit

    All other libraries used (os, shutil, pathlib, json) are part of the Python standard library — no extra installs required.


    File Category Mapping

    The organizer groups files by extension. Here is the complete mapping used in this project:

    CategoryExtensions
    Images.jpg .jpeg .png .gif .bmp .svg .webp .ico .tiff
    Documents.pdf .doc .docx .txt .xls .xlsx .ppt .pptx .odt .csv
    Videos.mp4 .mov .avi .mkv .wmv .flv .webm .m4v
    Audio.mp3 .wav .flac .aac .ogg .m4a .wma
    Archives.zip .tar .gz .rar .7z .bz2 .xz
    Code.py .js .ts .html .css .java .cpp .c .json .yaml .xml .sh
    Executables.exe .dmg .pkg .deb .rpm .msi .app
    Data.db .sqlite .sql .csv .parquet .feather
    OthersEverything else

    Project Structure

    code
    file_organizer/
    ├── file_organizer.py    ← Main Streamlit app
    ├── undo_log.json        ← Auto-generated undo history
    └── requirements.txt

    Step 1: Define the Category Map

    Create file_organizer.py and start with the extension-to-category mapping:

    code
    import os
    import shutil
    import json
    from pathlib import Path
    import streamlit as st
    
    st.set_page_config(page_title="File Organizer", page_icon="📂", layout="wide")
    st.title("📂 File Organizer")
    st.write("Automatically sort your files into categorised folders — preview first, move when ready.")
    
    CATEGORY_MAP = {
        "Images":      [".jpg", ".jpeg", ".png", ".gif", ".bmp", ".svg",
                       ".webp", ".ico", ".tiff"],
        "Documents":   [".pdf", ".doc", ".docx", ".txt", ".xls", ".xlsx",
                       ".ppt", ".pptx", ".odt", ".csv"],
        "Videos":      [".mp4", ".mov", ".avi", ".mkv", ".wmv", ".flv",
                       ".webm", ".m4v"],
        "Audio":       [".mp3", ".wav", ".flac", ".aac", ".ogg", ".m4a", ".wma"],
        "Archives":    [".zip", ".tar", ".gz", ".rar", ".7z", ".bz2", ".xz"],
        "Code":        [".py", ".js", ".ts", ".html", ".css", ".java", ".cpp",
                       ".c", ".json", ".yaml", ".xml", ".sh"],
        "Executables": [".exe", ".dmg", ".pkg", ".deb", ".rpm", ".msi", ".app"],
        "Data":        [".db", ".sqlite", ".sql", ".parquet", ".feather"],
    }
    
    def get_category(extension: str) -> str:
        ext = extension.lower()
        for category, extensions in CATEGORY_MAP.items():
            if ext in extensions:
                return category
        return "Others"

    Step 2: Scan the Folder

    The scanner collects all files (not folders) in the target directory and returns a list of dicts with metadata:

    code
    def scan_folder(folder_path: str) -> list[dict]:
        """Return a list of file info dicts for all files in the given folder."""
        path = Path(folder_path)
        files = []
        for item in path.iterdir():
            if item.is_file():
                size_bytes = item.stat().st_size
                files.append({
                    "name":      item.name,
                    "extension": item.suffix.lower() or "(none)",
                    "category":  get_category(item.suffix),
                    "size_kb":   round(size_bytes / 1024, 1),
                    "source":    str(item),
                    "dest":      str(path / get_category(item.suffix) / item.name),
                })
        return sorted(files, key=lambda f: f["category"])

    Note: item.iterdir() only lists the top level of the folder — it does not recurse into subfolders. Add rglob("*") instead of iterdir() if you want recursive scanning.


    Step 3: Build the Folder Input and Scan UI

    code
    st.divider()
    st.subheader("📁 Step 1 — Choose a Folder")
    
    folder_input = st.text_input(
        "Enter the full path to the folder you want to organise:",
        placeholder="e.g. /Users/you/Downloads  or  C:\\Users\\you\\Downloads",
    )
    
    dry_run = st.toggle("🔍 Dry Run (preview only — do not move files)", value=True)
    
    if st.button("🔎 Scan Folder", type="primary") and folder_input:
        if not os.path.isdir(folder_input):
            st.error("❌ That path does not exist or is not a folder. Please check and try again.")
        else:
            files = scan_folder(folder_input)
            st.session_state["scanned_files"] = files
            st.session_state["target_folder"] = folder_input
            st.success(f"✅ Found **{len(files)} files** in `{folder_input}`")

    Step 4: Display the Preview Table

    After scanning, show a filterable preview of what will be moved:

    code
    if "scanned_files" in st.session_state:
        files = st.session_state["scanned_files"]
    
        st.divider()
        st.subheader("📋 Step 2 — Preview")
    
        # Category filter
        categories = ["All"] + sorted(set(f["category"] for f in files))
        selected_cat = st.selectbox("Filter by category:", categories)
    
        filtered = files if selected_cat == "All" else [
            f for f in files if f["category"] == selected_cat
        ]
    
        # Summary metrics
        col1, col2, col3 = st.columns(3)
        col1.metric("Total Files", len(files))
        col2.metric("Categories", len(set(f["category"] for f in files)))
        col3.metric(
            "Total Size",
            f"{sum(f['size_kb'] for f in files) / 1024:.1f} MB"
        )
    
        # Preview table
        st.dataframe(
            [{"File": f["name"], "Category": f["category"],
              "Extension": f["extension"], "Size (KB)": f["size_kb"]}
             for f in filtered],
            use_container_width=True,
            hide_index=True,
        )

    Step 5: Move Files

    The core organizer function creates subfolders and moves each file:

    code
    def organize_files(files: list[dict], dry_run: bool) -> list[dict]:
        """Move files to their category subfolders. Returns a log of moves made."""
        log = []
        for f in files:
            dest_dir = Path(f["dest"]).parent
            if not dry_run:
                dest_dir.mkdir(parents=True, exist_ok=True)
                shutil.move(f["source"], f["dest"])
            log.append({
                "file":     f["name"],
                "from":     f["source"],
                "to":       f["dest"],
                "dry_run":  dry_run,
            })
        return log

    Connect this to the UI:

    code
    st.divider()
        st.subheader("🚀 Step 3 — Organise")
    
        action_label = "🔍 Simulate (Dry Run)" if dry_run else "📁 Organise Files"
        if st.button(action_label, type="primary", use_container_width=True):
            log = organize_files(files, dry_run)
    
            # Save undo log to JSON
            if not dry_run:
                undo_path = Path(st.session_state["target_folder"]) / "_organizer_undo.json"
                with open(undo_path, "w") as f_json:
                    json.dump(log, f_json, indent=2)
                st.success(f"✅ Organised {len(log)} files! Undo log saved to `_organizer_undo.json`.")
            else:
                st.info(f"🔍 Dry run complete — {len(log)} files would be moved. Toggle off Dry Run to apply.")
    
            st.session_state["last_log"] = log

    Step 6: Undo Functionality

    The undo feature reads the JSON log and moves each file back to its original location:

    code
    def undo_organize(log: list[dict]) -> int:
        """Move files back to their original locations using the undo log."""
        count = 0
        for entry in log:
            src = Path(entry["to"])
            dst = Path(entry["from"])
            if src.exists():
                dst.parent.mkdir(parents=True, exist_ok=True)
                shutil.move(str(src), str(dst))
                count += 1
        return count

    Add this to the UI after the organise section:

    code
    st.divider()
    st.subheader("↩️ Undo")
    
    undo_file = st.file_uploader(
        "Upload an `_organizer_undo.json` file to revert a previous organisation:",
        type=["json"],
    )
    
    if undo_file and st.button("↩️ Undo Last Organisation"):
        log_data = json.load(undo_file)
        moved_back = undo_organize(log_data)
        st.success(f"✅ Reverted {moved_back} files to their original locations.")

    Complete file_organizer.py

    code
    import os, shutil, json
    from pathlib import Path
    import streamlit as st
    
    st.set_page_config(page_title="File Organizer", page_icon="📂", layout="wide")
    st.title("📂 File Organizer")
    
    CATEGORY_MAP = {
        "Images":      [".jpg",".jpeg",".png",".gif",".bmp",".svg",".webp",".ico",".tiff"],
        "Documents":   [".pdf",".doc",".docx",".txt",".xls",".xlsx",".ppt",".pptx",".odt"],
        "Videos":      [".mp4",".mov",".avi",".mkv",".wmv",".flv",".webm",".m4v"],
        "Audio":       [".mp3",".wav",".flac",".aac",".ogg",".m4a",".wma"],
        "Archives":    [".zip",".tar",".gz",".rar",".7z",".bz2",".xz"],
        "Code":        [".py",".js",".ts",".html",".css",".java",".cpp",".c",".json",".yaml"],
        "Executables": [".exe",".dmg",".pkg",".deb",".rpm",".msi",".app"],
    }
    
    def get_category(ext): 
        ext = ext.lower()
        for cat, exts in CATEGORY_MAP.items():
            if ext in exts: return cat
        return "Others"
    
    def scan_folder(folder):
        p = Path(folder)
        return sorted([
            {"name": i.name, "extension": i.suffix.lower() or "(none)",
             "category": get_category(i.suffix),
             "size_kb": round(i.stat().st_size/1024, 1),
             "source": str(i), "dest": str(p/get_category(i.suffix)/i.name)}
            for i in p.iterdir() if i.is_file()
        ], key=lambda f: f["category"])
    
    def organize_files(files, dry_run):
        log = []
        for f in files:
            dest_dir = Path(f["dest"]).parent
            if not dry_run:
                dest_dir.mkdir(parents=True, exist_ok=True)
                shutil.move(f["source"], f["dest"])
            log.append({"file": f["name"], "from": f["source"], "to": f["dest"]})
        return log
    
    def undo_organize(log):
        count = 0
        for e in log:
            src, dst = Path(e["to"]), Path(e["from"])
            if src.exists():
                dst.parent.mkdir(parents=True, exist_ok=True)
                shutil.move(str(src), str(dst))
                count += 1
        return count
    
    # --- UI ---
    folder_input = st.text_input("📁 Folder path:", placeholder="/Users/you/Downloads")
    dry_run = st.toggle("🔍 Dry Run", value=True)
    
    if st.button("🔎 Scan", type="primary") and folder_input:
        if not os.path.isdir(folder_input):
            st.error("Path not found.")
        else:
            files = scan_folder(folder_input)
            st.session_state.update({"files": files, "folder": folder_input})
            st.success(f"Found {len(files)} files.")
    
    if "files" in st.session_state:
        files = st.session_state["files"]
        cats = ["All"] + sorted(set(f["category"] for f in files))
        sel = st.selectbox("Filter:", cats)
        shown = files if sel == "All" else [f for f in files if f["category"]==sel]
        c1, c2, c3 = st.columns(3)
        c1.metric("Files", len(files))
        c2.metric("Categories", len(cats)-1)
        c3.metric("Size", f"{sum(f['size_kb'] for f in files)/1024:.1f} MB")
        st.dataframe([{"File": f["name"], "Category": f["category"],
                       "Size (KB)": f["size_kb"]} for f in shown],
                     use_container_width=True, hide_index=True)
        label = "🔍 Simulate" if dry_run else "📁 Organise Files"
        if st.button(label, type="primary", use_container_width=True):
            log = organize_files(files, dry_run)
            if not dry_run:
                undo_path = Path(st.session_state["folder"]) / "_organizer_undo.json"
                with open(undo_path, "w") as fj: json.dump(log, fj, indent=2)
                st.success(f"✅ Organised {len(log)} files!")
            else:
                st.info(f"🔍 {len(log)} files would be moved.")
    
    st.divider()
    undo_file = st.file_uploader("Upload undo JSON", type=["json"])
    if undo_file and st.button("↩️ Undo"):
        st.success(f"Reverted {undo_organize(json.load(undo_file))} files.")

    Run the App

    code
    streamlit run file_organizer.py

    Key Python Standard Library Tools

    ModuleFunction / ClassWhat It Does
    pathlibPath(folder)Cross-platform file paths
    pathlibpath.iterdir()Lists contents of a directory
    pathlibpath.is_file()True only for files (not dirs)
    pathlibpath.stat().st_sizeFile size in bytes
    pathlibpath.suffixFile extension, e.g. .pdf
    pathlibpath.mkdir(parents=True)Creates folder (and parents)
    shutilshutil.move(src, dst)Moves a file to a new location
    jsonjson.dump() / json.load()Saves and loads the undo log
    osos.path.isdir()Validates that a path is a folder

    Comparing Approaches: shutil vs pathlib rename

    When moving files in Python you have two main options:

    ApproachCodeCross-device?Notes
    shutil.move()shutil.move(src, dst)✅ YesWorks across drives; preferred
    Path.rename()Path(src).rename(dst)❌ NoFast but only same filesystem
    Path.replace()Path(src).replace(dst)❌ NoOverwrites destination silently

    Always use shutil.move() when the destination might be on a different drive or network share.


    Safety Best Practices

    RiskMitigation in This App
    Accidentally moving important filesDry run toggle on by default
    Overwriting existing files at destinationUnique filenames are preserved; add a counter suffix if needed
    Moving already-organised subfolderspath.is_file() skips directories
    No way to recoverUndo log saved as JSON before every operation
    Path doesn't existos.path.isdir() check before scanning

    Extending the File Organizer

    FeatureHow to Implement
    Recursive scanningReplace iterdir() with rglob("*")
    Custom category rulesAdd a UI form to define custom extension→category mappings
    Duplicate detectionHash files with hashlib.md5() and flag duplicates
    Schedule automatic runsUse schedule library or a cron job
    Email reportSend the move summary via smtplib
    Dark/light mode UIAlready handled by Streamlit's built-in theme switcher

  • Password Generator — Another useful Python automation tool built with Streamlit.
  • Weather App — Uses requests and external APIs to automate data retrieval.
  • Interactive Quiz App — Explores st.session_state for multi-step workflows, the same pattern used in this organizer.

  • Conclusion

    You have built a fully functional, safe, and reversible File Organizer in Python and Streamlit. The app uses only the Python standard library (plus Streamlit) — no heavy dependencies, no cloud services, no API keys. The dry-run mode and undo log make it production-safe enough to run on your real Downloads folder.

    Resources:

  • Python pathlib Documentation
  • Python shutil Documentation
  • Streamlit st.dataframe
  • Streamlit st.toggle
  • Adapted from: File Organizer using Python and Streamlit

    Checkpoint
    The organizer sorts a messy folder by file type in one run.
    What you learned
    • pathlib and shutil in practice
    • Comprehensions filtering real files
    • Standard library confidence