Enter your licence key to continue. The same key works on every device you own.
Local LLM Studio runs models on your own machine. These are the pieces it needs.
Pick a .gguf above and press Import into Ollama.
No Forge log yet — start Forge to see its output here.
Use Gallery in the header (or Open Gallery) to browse Studio + local Forge outputs inside this app. Forge IIB opens the full extension in a new tab.
What the tagging model was asked, the raw line it answered, and what the server dropped on the way to Forge. A window on data already passing through this page — it adds no model work and does not slow generation. Kept for the last 12 runs, cleared on reload.
Second pass at a larger size. This is how detail is actually added — generating straight to a big canvas duplicates subjects instead. Measured on an 8 GB card at 896×1152: 1.25× costs about 3× the time, 1.5× about 5×. 2× runs out of VRAM and takes Forge down with it, so it is capped below that.
Large side avatars + chat wallpaper. Point each at a folder of images. Cycle 5–60s or static. Chat: /avatar me C:\pics · /avatar model … · /bg …
Defaults for new bubbles. Use Folder (cycle) or Image (static). Click any bubble avatar for per-bubble settings.
The longest a single reply may be. On a long chat the context window below usually runs out first.
How much of the conversation the model can see, in tokens. Blank uses a safe default (8,192) chosen for the largest model this ships against. Raising it lets long chats keep going — and costs VRAM, so if image generation starts spilling to system memory, lower it again. When a chat finally outgrows this, Ollama refuses the request rather than dropping the oldest messages, so you get a clear error instead of a model that has quietly forgotten the character. Raise it, or start a new chat.
Two problems sampling cannot fix: the model writing your turn for you, and the model forgetting who it is after twenty exchanges.
Re-inserted near the end of the conversation on every turn, not at the top. The system prompt is thousands of tokens behind by turn 30 and the model stops weighting it — this sits close to where the reply is written, so it cannot be crowded out. Keep it to two or three lines.
How many messages back from the end it goes. 0 places it after your message — strongest, but the model may answer it instead of the story. 3 is the usual choice.
Generation halts if the model produces one. One per line. Each is matched at the start of a line, so type Moodm: to stop the model speaking as you — the line break is added for you, which is what keeps it from matching the same word mid-sentence.
Speaker labels such as User: at the start of a line, and leaked ChatML control tokens. The console shows the armed list under stops. If replies ever end early for no clear reason, clear this box first — Ollama reports a stop-sequence hit and a natural ending identically, so this switch is the only way to tell them apart.
Blank means leave it to the model. Every value here shows in the chat console under effective, so you can confirm it took.
Fix it and the same prompt gives the same reply, so a change you make is attributable to the setting rather than to chance — as long as the model stays loaded exactly as it is. Measured 2026-08-17: back-to-back runs reproduce whether the model is fully on the GPU or partly on the CPU, but anything that makes Ollama RELOAD it — switching model, changing Context window, leaving it idle until it unloads — can change the answer for the same seed, because the layers are split differently the next time and the arithmetic lands in a different order. So do your A/B runs back to back, and re-take the baseline after any model switch.
Scales the cutoff to the model’s confidence instead of a fixed probability mass, so it stays coherent at the higher temperatures creative writing wants. Try 0.05 with Top P relaxed.
How far back the penalty can see, in tokens. The default 64 is why a paragraph repeated 400 tokens ago survives it. For long stories try 256–512.
Gentler than the repeat penalty, which at high values starts damaging grammar — the model avoids common words because they keep recurring.
Tokens pinned at the front — where the system prompt and the character card live — so a very long reply cannot push them out as it runs. It is not what protects you from a full window. Measured 2026-08-17: this Ollama does not trim an over-long conversation, it refuses it, with request (1078 tokens) exceeds the available context. So a chat that outgrows the window stops with an error rather than quietly forgetting who the character is — raise Context window, or start a new chat.
"Tokens: 7,948 in · 16 out". The IN count is the only warning that the context window is filling up, and the OUT count is what says a reply was cut off rather than finished.
No chat turns yet. Send a message.
Logs every turn as it starts and again when it finishes — the options actually sent, the tokens in and out, and which limit stopped a reply. The bubble can only show the outcome; this shows the cause.
Applies to local and external chat. Also sets Max Tokens unless you change tokens after saving.
Real Google web + image results via
Serper.dev.
Free tier: sign up → API key. Use chat buttons or
/web … / /images ….
Local Kokoro-82M voices. Click Speak on any assistant reply, or enable auto-play.
Starts reading after the first sentence instead of waiting for the whole reply. Same as the 🔊 icon on a reply.
Kokoro voice files — .pt,
.npy, .npz or a whole
.bin pack. Imported voices live in
their own file, so re-downloading the models
never deletes them.
Piper (.onnx) and RVC (.pth) are
different engines and will be refused.
Welcome to Local LLM Studio! Select and load a model to start chatting.
New features:
/avatar me C:\pics · /avatar model … · /bg … set defaults/imagine … or the image button for Forge generationUses the Forge extension stable-diffusion-webui-wd14-tagger (local, needs --api). Drop an image, pick a gallery file, or tag the last chat image.
Images always open in a new bubble under the last reply (model text is kept).
2girls and
solo is dropped.
Without regions, traits bleed between
them: a plain prompt is one bag of words,
so the model cannot tell which hair colour belongs
to which face. Regions condition each character
separately and stop that.
Second pass at a larger size. This is how detail is actually added — generating straight to a big canvas duplicates subjects instead. Measured on an 8 GB card at 896×1152: 1.25× costs about 3× the time, 1.5× about 5×. 2× runs out of VRAM and takes Forge down with it, so it is capped below that.
Select a profile on the left, edit fields, then Save Changes. Your old prompts are in My Current Setup. Use Strict No-Refuse for the hard compliance profile.
How a turn is written — pacing, point of view, how much to spend at once. The profile above still decides what may be written; a style only shapes the form. Pick one per character card, under Scenario style.
This is you — the person the character is speaking to. It is stated to the model after the character, so the scene has two described sides instead of one.
Each entry joins the prompt only when one of its keywords appears in the last few messages. Facts about the world — places, factions, history. Who the character is belongs on the character card; who you are belongs on your persona.
Paste a JSON object (or contents of a .json / .txt file) below, then Apply.
Are you sure you want to delete this chat? This action cannot be undone.