The Column Orders.

Portico doesn't lock you into one model. Point it at any open-weight model, and it handles the rest. To make it easy to choose, we've categorized our recommended models into architectural orders based on how much RAM your machine has.

Open Weight Models Auto-sized to hardware Drag & Drop custom models

Don't see your model on this list? No problem. Portico supports any standard GGUF model file. Just drag it into the app, and it will run.


How much memory do you have?

That single number decides more than anything else about which models you can run. Find your row, then read the order below it.

Named after the classical orders, in ascending order. A model larger than your memory will still load — it just reads from disk on every word, which is far too slow to use. Portico checks your machine and marks what fits.


Doric — about 3B, 8 GB RAMQuick drafts, older laptops.

Runs fast on your card (~50 tok/s). The best choice for quick questions, summarizing text, or running on machines that weren't built for heavy lifting.

Llama 3.2 1B 1.3 GB

Fastest here; quick questions on any machine.

Llama 3.2 3B 2.0 GB

The balanced default for everyday tasks.

Gemma 2 2B 1.6 GB

Google's lightweight model; surprisingly sharp.

Phi 3.5 Mini 2.4 GB

Reasoning above its weight class.

Stable LM 2 1.6B 1.5 GB

Ultra-light and fast for basic text generation.

Qwen2.5-VL 3B + projector 3.2 GB

Sees and understands images.


Ionic — about 8B, 16 GB RAMYour machine's sweet spot.

The best balance of quality and speed for general reasoning and code. If you have a modern laptop with 16 GB of RAM, this is where you should live.

Qwen 2.5 7B 4.7 GB

Clear step up from 3B class models.

Llama 3.1 8B 4.9 GB

All-rounder with massive 128K context.

Llama 3 8B 4.7 GB

The previous generation; still excellent for coding.

Qwen3 8B 5.0 GB

Newer generation, thinks before answering.

Mistral 7B v0.3 4.4 GB

The classic workhorse. Fast and reliable.

Gemma 2 9B 5.8 GB

Google's model; the most natural writing style.

Mistral Nemo 12B 7.5 GB

128K context, excellent for creative writing.

DeepSeek R1 Distill 7B 4.7 GB

Shows its reasoning steps openly.

Qwen 2.5 Coder 7B 4.7 GB

Specialized for programming and code generation.

StarCoder 2 7B 3.7 GB

Code completion and generation across many languages.

Qwen 2.5 Math 7B 4.7 GB

Maths only — a great study aid.

Aya 23 8B 5.1 GB

Excellent Spanish + 22 other languages.

Hermes 3 8B 4.9 GB

Follows complex instructions and personas.


Corinthian — 14–32B, 32 GB RAMDesktops and workstations.

Slow on your laptop but highly capable. These models are designed for dedicated hardware and offer massive leaps in reasoning, coding, and prose.

DeepSeek R1 Distill 14B 9.0 GB

Strongest reasoning you can run locally.

Qwen 2.5 Coder 14B 9.0 GB

A serious coding assistant for complex repos.

Qwen3 14B 9.0 GB

Best quality-per-GB with a thinking mode.

Phi 3 Medium 14B 8.2 GB

Microsoft's mid-tier; excels at logic and reasoning.

Gemma 2 27B 16.7 GB

Beautiful prose — needs ~19 GB RAM to run well.

Yi 34B 19.0 GB

Strong multilingual capabilities and deep context.

Command R 35B 20.0 GB

Built specifically for RAG and tool use.

Mixtral 8x7B 24.0 GB

Mixture-of-Experts; fast responses, high quality.

Qwen3 32B 19.8 GB

Near-flagship performance locally.

DeepSeek R1 Distill 32B 19.9 GB

Deepest reasoning capabilities for complex logic.


Composite — 70B, 64 GB RAMThe flagship tier.

For high-end workstations with massive amounts of memory. These models offer frontier-class intelligence, but are completely out of reach on standard laptops.

Llama 3.3 70B 42.5 GB

Flagship-class. Out of reach on this laptop, but stunning on a workstation.

Qwen 2.5 72B 42.0 GB

Top-tier open model. Exceptional at coding and math.

Mixtral 8x22B 42.5 GB

Massive MoE architecture. Incredible reasoning.

Images and speech.

These don't fit the column metaphor since they aren't chat models, but Portico supports them natively for multimedia creation and transcription.

🎨 Image Generation (measured on your GPU)

DreamShaper 8 2.1 GB

15s generation — best value, your default.

Realistic Vision 6 2.1 GB

15s generation — highly photographic.

SD 1.5 base 2.1 GB

15s generation — the plain original model.

LCM DreamShaper 2.1 GB

Ultra-fast (2s) generation for instant drafts.

SDXL Turbo 6.9 GB

84s generation @ 768px resolution.

Juggernaut XL v9 7.1 GB

197s @ 1024px — the absolute best quality.

SDXL base 1.0 6.9 GB

Slow generation; stronger prompt following.

FLUX.1 [schnell] 11.5 GB

Stunning prompt adherence and text generation.

🎙️ Speech

Whisper Base 150 MB

Bundled engine. Supports ~99 languages for local transcription.

Whisper Large v3 3.1 GB

Maximum accuracy for complex audio and accents.

Piper TTS ~200 MB

Ultra-fast local text-to-speech in dozens of voices.

Try it on your own machine.

Free and open source. Works with the open models you already trust.