Skip to content

Repository files navigation


NoteForge

Turn a public course video into structured Markdown study notes.

Python 3.11+ CLI Tests

Bilibili subtitle collection, LLM-powered knowledge extraction, and traceable notes

Quick start · Configuration · Usage · Troubleshooting · Development

English · 简体中文


Warning

NoteForge is still under active development and has many incomplete or unpolished areas. Use it with caution. The author assumes no responsibility for any loss or damage resulting from its use.

NoteForge CLI example


Why NoteForge

Long course videos are useful, but turning them into reviewable notes takes time. NoteForge collects a video's available subtitles, identifies semantic sections, extracts key concepts with an LLM, and writes a structured Markdown document with source timestamps.

Bilibili URL
    ↓
Subtitle discovery, selection, download, and normalization
    ↓
Transcript chunking and semantic analysis
    ↓
Knowledge-point extraction
    ↓
Structured Markdown notes with timestamps

NoteForge downloads subtitles only. It does not download the video or audio.


Capabilities

Capability What it does
Source inspection Normalizes a URL and shows video, subtitle, and transcript metadata
Subtitle selection Prefers a requested language, then supported Chinese and English tracks
LLM analysis Supports OpenAI-compatible, Anthropic Messages, and Ollama API formats
Note generation Creates organized Markdown notes with concepts, explanations, and timestamps
Multi-part videos Handles the p parameter in Bilibili multi-part video URLs

Current scope: end-to-end note generation supports standard Bilibili video URLs with an available VTT or SRT subtitle track.


Requirements

  • Python 3.11 or newer
  • uv
  • Network access to Bilibili and the selected model endpoint
  • A supported subtitle track on the target video
  • For restricted videos: a locally installed, signed-in browser

Ollama users also need a running Ollama installation and a local model that follows JSON-output instructions reliably.


Quick start

1. Install the command

uv tool install noteforge-cli

Confirm the CLI is ready:

noteforge --version
noteforge --help

If your shell cannot find the command after installation, run uv tool update-shell, restart the terminal, and try again.

2. Configure an LLM

Start the interactive setup:

noteforge configure

For a first-run check, verify the configured model and optionally inspect whether a video needs browser cookies and has supported subtitles:

noteforge doctor \
  "https://www.bilibili.com/video/BVxxxxxxxxxx"

For the default local Ollama setup, accept ollama, then prepare the suggested model:

ollama pull qwen2.5:7b

The wizard stores the settings in a local .env file with user-only permissions. NoteForge loads this file automatically. The default setup expects Ollama at http://localhost:11434; keep ollama serve running if your installation does not start it automatically.

For an OpenAI-compatible or Anthropic Messages endpoint, choose its API format in the wizard. API-key input is hidden. You may also copy and edit .env.example; see LLM configuration.

3. Inspect a video before generating

noteforge inspect \
  "https://www.bilibili.com/video/BVxxxxxxxxxx" \
  --cookies-from-browser chrome

Check that the JSON output contains a non-null selected_subtitle and transcript, and that segment_count is greater than zero.

4. Generate notes

noteforge generate \
  "https://www.bilibili.com/video/BVxxxxxxxxxx" \
  --output output/note.md \
  --cookies-from-browser chrome

The completed note is written to output/note.md. Parent directories are created automatically.


LLM configuration

NoteForge reads configuration from a .env file in the current directory and from process environment variables. Values explicitly exported in the shell take precedence over .env.

NOTEFORGE_LLM_PROVIDER is a legacy-compatible 0.1 name. Its value selects the API wire format, not the company operating the model endpoint. For example, a compatible DeepSeek endpoint uses the openai format with its own BASE_URL.

Run the setup again whenever you want to change API formats, endpoints, or models:

noteforge configure

When generate detects missing configuration in an interactive terminal, it offers to start the same wizard automatically. In scripts and CI it exits with a clear instruction instead of waiting for input.

Variable Required Description
NOTEFORGE_LLM_PROVIDER Yes API format identifier: ollama, openai, or anthropic; the legacy variable name is retained for 0.1 compatibility
NOTEFORGE_LLM_MODEL Yes Model identifier accepted by the target endpoint
NOTEFORGE_LLM_API_KEY OpenAI-compatible/Anthropic Messages Target endpoint API key; not needed by Ollama
NOTEFORGE_LLM_BASE_URL No Model endpoint; each API format has an official default and can target compatible services
NOTEFORGE_LLM_TIMEOUT_SECONDS No Request timeout in seconds; default: 60

Ollama

NOTEFORGE_LLM_PROVIDER=ollama
NOTEFORGE_LLM_MODEL=qwen2.5:7b
NOTEFORGE_LLM_BASE_URL=http://localhost:11434
NOTEFORGE_LLM_TIMEOUT_SECONDS=120

OpenAI-compatible

NOTEFORGE_LLM_PROVIDER=openai
NOTEFORGE_LLM_MODEL=<an-available-chat-completions-model>
NOTEFORGE_LLM_API_KEY=<your-api-key>
NOTEFORGE_LLM_TIMEOUT_SECONDS=120

This format uses the OpenAI Chat Completions request and response schema. Its default endpoint is https://api.openai.com/v1; set NOTEFORGE_LLM_BASE_URL to use DeepSeek or another compatible endpoint. Selecting openai does not require OpenAI to be the provider.

Anthropic Messages

NOTEFORGE_LLM_PROVIDER=anthropic
NOTEFORGE_LLM_MODEL=<an-available-anthropic-model>
NOTEFORGE_LLM_API_KEY=<your-api-key>
NOTEFORGE_LLM_TIMEOUT_SECONDS=120

This format uses the Anthropic Messages request and response schema. Its default endpoint is https://api.anthropic.com/v1; compatible endpoints can be configured as well.

Never commit .env or a real API key. .env is ignored by Git.


Usage

Inspect

Use inspect to validate collection and subtitle access without calling an LLM:

noteforge inspect VIDEO_URL [OPTIONS]

Useful options:

--cookies-from-browser TEXT  Browser used for cookies: chrome, edge, firefox, safari
--subtitle-language TEXT     Preferred language, for example zh-Hans, zh-CN, or en
--subtitle-output-dir PATH   Subtitle cache root

To try a public video without browser cookies:

noteforge inspect VIDEO_URL --cookies-from-browser ""

Generate

noteforge generate VIDEO_URL [OPTIONS]

Examples:

# Prefer Simplified Chinese subtitles
noteforge generate VIDEO_URL \
  --subtitle-language zh-Hans \
  --output output/course-note.md

# Generate notes for part 2 of a multi-part video
noteforge generate \
  "https://www.bilibili.com/video/BVxxxxxxxxxx?p=2" \
  --output output/part-2.md

# Do not read browser cookies
noteforge generate VIDEO_URL \
  --cookies-from-browser "" \
  --output output/note.md

Run noteforge COMMAND --help for the complete option reference.

Run records

Every generate invocation creates .noteforge/runs/<run-id>/. Successful, failed, and cancelled runs are all retained with a manifest, JSONL event stream, stage artifacts, final-note copy, and structured error details. Use --run-dir PATH to choose another records root.

Run records never persist API keys, cookies, or Authorization headers. Model endpoints are stored only as origins with credentials, paths, and query strings removed.


Troubleshooting

缺少配置:NOTEFORGE_LLM_PROVIDER

Run the interactive configuration:

noteforge configure

Browser-cookie errors

Use the name of a browser installed on this machine and make sure it has a signed-in Bilibili session:

noteforge inspect VIDEO_URL --cookies-from-browser firefox

For a public video, retry without cookies:

noteforge inspect VIDEO_URL --cookies-from-browser ""

Close the browser temporarily if its cookie database is locked.

HTTP 412 or platform risk control

Use cookies from a signed-in browser, avoid repeated rapid requests, and retry later. Platform-side risk control cannot be eliminated by NoteForge.

No supported subtitle

Confirm the video exposes a subtitle track in Bilibili, try a preferred language with --subtitle-language, and inspect the subtitle_tracks output. NoteForge currently parses VTT and SRT tracks; it does not transcribe audio.

The LLM returns invalid JSON

Use a model with strong instruction-following and structured-output ability. For a small local model, try a larger model or increase the timeout. The partial output is not written as a completed note.

Connection or timeout failures

Check the provider URL and API key, then increase:

NOTEFORGE_LLM_TIMEOUT_SECONDS=180

Development

git clone https://github.com/ztygod/NoteForge.git
cd NoteForge
uv sync --group dev
uv run pytest -q
uv build

Release installation, upgrade, and removal:

uv tool install noteforge-cli
uv tool upgrade noteforge-cli
uv tool uninstall noteforge-cli

Publishing a release

The noteforge name is already owned by another project on PyPI. This project is therefore distributed as noteforge-cli while continuing to expose the noteforge terminal command.

Before the first release, register a pending Trusted Publisher on PyPI with:

PyPI project name: noteforge-cli
GitHub owner: ztygod
GitHub repository: NoteForge
Workflow: publish.yml
Environment: pypi

Create a protected pypi environment in the GitHub repository and require manual approval. Then publish a version by updating the version and pushing a matching tag:

uv version 0.1.0
git add pyproject.toml uv.lock
git commit -m "release: v0.1.0"
git tag v0.1.0
git push origin main v0.1.0

The workflow builds and validates both distributions before publishing through PyPI Trusted Publishing; no long-lived PyPI token is stored in GitHub.

Project structure:

noteforge/
├── src/noteforge/
│   ├── cli/          # Typer commands
│   ├── collector/    # Source inspection and Bilibili collection
│   ├── subtitle/     # Subtitle selection, parsing, and normalization
│   ├── knowledge/    # Chunking, semantic analysis, and extraction
│   ├── llm/          # OpenAI-compatible, Anthropic Messages, and Ollama adapters
│   ├── document/     # Learning-document construction
│   ├── renderer/     # Markdown rendering and writing
│   └── core/         # End-to-end pipeline
├── tests/
├── .github/workflows/publish.yml
├── .env.example
├── pyproject.toml
└── uv.lock

From a long video to notes you can actually review.

About

An AI-powered knowledge extraction tool. Transform lectures, videos and documents into structured Markdown notes.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages