Files
AndyTranscribe/README.md
T
ben e0400aadef Add live transcription progress and Docker Whisper.
Surface stage, percent, and elapsed time while jobs run, and ship Compose so local confidential transcription can use faster-whisper on port 8090.
2026-08-12 14:32:56 +02:00

149 lines
4.4 KiB
Markdown

# AndyTranscribe
Upload pocket-recorder audio (MP3, WAV, OGG, and more), extract embedded metadata, and transcribe with OpenAI Whisper, a local faster-whisper server, or a remote OpenAI-compatible endpoint.
Built with Laravel 13, Blade, Tailwind CSS 4, and [Laravel AI](https://github.com/laravel/ai).
## Features
- Upload common audio formats (MP3, WAV, OGG, FLAC, M4A, AAC, WebM, WMA, AIFF — up to 100 MB)
- Automatic metadata extraction when tags are present (title, artist, album, duration, recorded date)
- Search recordings by title, artist, or transcript
- Queued transcription with three engines:
- **Cloud** — OpenAI Whisper (`whisper-1`)
- **Local** — confidential; OpenAI-compatible [faster-whisper-server](https://github.com/fedirz/faster-whisper-server) via Docker
- **Ollama host** — user-supplied host URL exposing `/v1/audio/transcriptions`
- Live transcription progress (stage, %, elapsed time)
- Copy finished transcripts from the recording detail page
## Requirements
- PHP 8.3+ (8.5 recommended)
- Composer
- Node.js & npm
- SQLite (default) or another supported database
- For **cloud** transcription: an OpenAI API key
- For **local** transcription: [Docker](https://docs.docker.com/get-docker/) (runs Whisper in a container)
- For **remote** transcription: a host with an OpenAI-compatible transcription API
## Setup
```bash
composer setup
```
That installs PHP and JS dependencies, copies `.env` if needed, generates the app key, runs migrations, and builds frontend assets.
Or step by step:
```bash
composer install
cp .env.example .env
php artisan key:generate
touch database/database.sqlite # if using SQLite
php artisan migrate
npm install
npm run build
```
## Local Whisper (Docker)
The **Local** engine does not run Whisper inside PHP. It calls an OpenAI-compatible HTTP API. This project ships Compose for that:
```bash
# CPU (works everywhere; slower on long files)
docker compose up -d whisper
# Optional: NVIDIA GPU
docker compose --profile gpu up -d whisper-gpu
```
First start downloads the model into a Docker volume (can take a few minutes).
Check it:
```bash
curl -s http://127.0.0.1:8090/health
```
Laravel talks to it at `LOCAL_WHISPER_URL` (default `http://127.0.0.1:8090/v1`). Port **8090** is used so it does not conflict with `php artisan serve` on 8000.
Stop:
```bash
docker compose down
```
Without this container, **Cloud** and **Ollama host** still work; only **Local** needs Docker.
## Configuration
Copy values from `.env.example`. The transcription-related settings are:
| Variable | Purpose |
| --- | --- |
| `OPENAI_API_KEY` | Required for cloud Whisper |
| `OPENAI_URL` | OpenAI API base URL (default `https://api.openai.com/v1`) |
| `LOCAL_WHISPER_URL` | Local faster-whisper base URL (default `http://127.0.0.1:8090/v1`) |
| `LOCAL_WHISPER_API_KEY` | API key for local server (often unused) |
| `LOCAL_WHISPER_MODEL` | Model name for local transcription |
| `REMOTE_WHISPER_MODEL` | Model name for Ollama-host transcription |
| `WHISPER_HOST_PORT` | Host port published by Compose (default `8090`) |
| `TRANSCRIPTION_TIMEOUT` | Job/HTTP timeout in seconds (default `600`) |
| `QUEUE_CONNECTION` | Use `database` (default) so transcription runs in the background |
Ensure `APP_URL` matches how you access the app (default `http://localhost:8000`).
## Running locally
Start the app, queue worker, and Vite together:
```bash
composer run dev
```
For confidential local transcription, also start Whisper:
```bash
docker compose up -d whisper
```
Or separately:
```bash
php artisan serve
php artisan queue:work
npm run dev
```
Open [http://localhost:8000/recordings](http://localhost:8000/recordings).
Transcription jobs are queued — keep a queue worker running or jobs will stay pending.
## Usage
1. **Upload** audio from Recordings → Upload (optional title override).
2. Open the recording and choose a transcription engine.
3. Watch live progress on the recording page until the transcript appears.
4. Search the list by title, artist, or transcript text.
## Transcription engines
| Driver | When to use | Needs |
| --- | --- | --- |
| `cloud` | Fastest path; audio leaves your machine | `OPENAI_API_KEY` |
| `local` | Confidential; audio stays on this machine | `docker compose up -d whisper` |
| `ollama` | Another machine on your network | Host URL + OpenAI-compatible `/v1/audio/transcriptions` |
## Tests
```bash
composer test
# or
php artisan test
```
## License
MIT