⚠ Hoca.me şu an test yayını aşamasındadır.
← Home Docs

How to Use Hoca?

Turn your PDF, LaTeX, or comic book into a narrated educational video.

🎬 Step-by-Step Pipeline

Hoca converts your document into video in a few steps. Narration is written per slide, synthesized by your chosen voice engine, and assembled into a slideshow video.

1
Upload File Drag and drop a PDF, LaTeX (.tex), or CBZ/CBR comic book file.
2
Choose Language & Engine Select content language, voice engine, and LLM engine. 75 languages are supported.
3
Narration Generation The LLM generates narration text for each slide. You can choose style (normal/short/detailed/comic/verbatim).
4
Voice Synthesis Your chosen TTS engine synthesizes narration text to speech (Edge TTS, Piper, XTTS, Chatterbox, Azure, Kokoro).
5
Video Render Slides + audio are combined; download MP4 video or optional SCORM package.
You can close the page after submitting — you'll receive an email notification when done (enable in settings).

🔊 Voice Engines

Each engine has different language support, quality, and cost profile.

Edge TTS Default

Microsoft cloud voice. 75+ languages, fast, free. Requires internet.

Piper Local

Open source, offline, TR/EN and others. Slightly more robotic than Edge.

XTTS GPU

High-quality voice clone. Requires reference WAV. Very fast with GPU.

Chatterbox GPU

MIT licensed, 23 langs + TR. Turbo mode is EN-only with [laugh] cues.

Azure Neural

Microsoft Azure neural voice. Requires AZURE_SPEECH_KEY.

Kokoro Local

ONNX-based, lightweight, high quality EN/JP/TR/FR/ES.

🤖 Narration (LLM) Engines

The LLM writes narration text for each slide. More capable models produce more natural narration.

Groq Default

Fast inference, Llama/Qwen models. Sufficient for most use cases.

Anthropic Claude

Highest quality narration, required for comic mode (vision). 2× credit multiplier.

Ollama Local

Fully local, no internet required. Performance depends on hardware.

🎭 Narration Styles

Normal

Default. Academic/professional, ~60-120 words/slide.

Short

~30-60 words. Ideal for quick overviews.

Detailed

~140-220 words. Deep explanation, for lecture content.

Comic Anthropic

Comic book mode. Speaker detection and multi-voice narration via Claude vision.

Verbatim

Reads page text verbatim via OCR. Supports karaoke highlighting.

Expressive Chatterbox Turbo

Narration with emotional cues like [laugh], [sigh]. English only.

📦 Export Formats

MP4

Standard video. Plays on any platform.

SCORM

ZIP package for LMSes (Moodle, Canvas). Includes completion tracking.

Audiobook Bundle

ZIP with MP3 + slide images + JSON. Opens in reader.hoca.me.

Audiobook

Single MP3 file. Combines narration from all slides.

Check the SCORM box in the Advanced tab to generate a SCORM package.

📖 Audiobook & Read-Along

Audiobook

Enable MP3 or Bundle output. Open the bundle in reader.hoca.me for interactive navigation.

Read-Along

Each sentence is highlighted on the PDF slide as audio plays. Check 'Read-Along' in the Audio tab.

Karaoke (Verbatim)

In Verbatim mode, the 'Karaoke' option highlights each word on the slide in yellow.

📱 reader.hoca.me

Open Hoca's bundles in your browser or as a PWA (installable app).

1
Download Bundle When your Hoca job completes, download the ZIP with 'Download Bundle'.
2
Open in Reader Open reader.hoca.me and drag the ZIP or click 'Upload File'.
3
Listen & Follow Audio starts automatically. Text follows in read-along mode. Use arrow keys to navigate slides.

PWA Installation

Click the 'Install' icon in your browser's address bar to run the reader as a desktop/mobile app. Works offline.

→ Reader's own help page (reader.hoca.me/docs)