← From Zero to Deployed Alle LektionenAll lessons Setup ModelleModels

Modelle & Benchmarks · Orientierung, kein Ranking-WettlaufModels & benchmarks · orientation, not a ranking race

Welches Modell nehme ich?Which model do I pick?

Ein Leitfaden plus ein Momentaufnahme-Überblick der Modellklassen. Ranglisten ändern sich wöchentlich — die Strategie nicht.A guide plus a snapshot overview of the model classes. Leaderboards change weekly — the strategy doesn't.

Block ABlock ADer LeitfadenThe guide

Generische Probleme

Generic problems

Nimm ein aktuelles Frontier-ModellTake a current frontier model

Für die allermeisten Aufgaben: ein aktuelles Frontier-Modell eines großen Labors (OpenAI, Anthropic, Google, xAI). Lauf nicht jede Woche dem Platz-1-Modell hinterher. Stabilität, gutes Tool-Use, langer Kontext und ein guter Agent drumherum schlagen zwei Punkte auf einem Leaderboard.

For the vast majority of tasks: a current frontier model from a big lab (OpenAI, Anthropic, Google, xAI). Don't chase the rank-1 model every week. Stability, solid tool-use, long context and a good agent around it beat two points on a leaderboard.

Sehr spezifische Domänen

Very specific domains

Frontier zuerst — Spezialmodell nur mit BelegFrontier first — specialist only with proof

Für enge Domänen (Medizin, Bio, Recht) startest du trotzdem mit einem Frontier-Modell als Default. Dann prüfst du, ob ein dediziertes oder fein-getuntes Modell nachweislich besser ist — über MedQA, PubMedQA, Medmarks oder deine eigenen Testdaten. Nimm das Spezialmodell nur, wenn Benchmark und dein Use Case es stützen — nicht, weil „Med“ im Namen steht.

For narrow domains (medicine, bio, law) you still start with a frontier model as the default. Then you check whether a dedicated or fine-tuned model is provably better — via MedQA, PubMedQA, Medmarks or your own test data. Pick the specialist only if the benchmark and your use case support it — not because the name contains "Med".

Immer Quelle und Datum mitzeigen. Ein Benchmark-Wert ohne „woher“ und „wann“ ist wertlos. Und: Leaderboards sind ein Filter, keine Wahrheit — sie messen enge Aufgaben unter idealen Bedingungen, nicht deinen Alltag.

Always show source and date. A benchmark number without "where from" and "when" is worthless. And: leaderboards are a filter, not truth — they measure narrow tasks under ideal conditions, not your day-to-day.

Die ModellklassenThe model classes

Frontier (proprietär)Frontier (closed)

GPT · Claude · Gemini · Grok

Default für die meisten Aufgaben.

Default for most work.

Frontier (offene Gewichte)Frontier (open-weight)

Llama · Qwen · DeepSeek · Mistral · GLM · Kimi

Self-Hosting, Datenkontrolle, Anpassung.

Self-hosting, data control, customisation.

Coding / AgentenCoding / agents

Codex · Qwen Coder · Devstral · Codestral

Code schreiben, Tools nutzen, Agenten.

Writing code, using tools, agents.

Reasoning / STEMReasoning / STEM

o-Serie · R1 · „Thinking“-Varianten

Mehrschritt-Logik, Mathe, Beweise.

Multi-step logic, math, proofs.

Günstig / schnellCheap / fast

mini · flash · haiku · nano · lite

Hohes Volumen, niedrige Latenz, wenig Kosten.

High volume, low latency, low cost.

Domäne: MedizinDomain: medical

Bio-Medical-Llama · II-Medical · …

Fein-getunt — nur mit Beleg statt als Default.

Fine-tuned — only with proof, not as a default.

Bild/Video (Imagen, Veo, Sora, Lyria) sind eine eigene Welt und hier bewusst ausgeklammert.Image/video (Imagen, Veo, Sora, Lyria) are their own world and deliberately left out here.

Wichtig (Medizin/Recht): Nichts auf dieser Seite ist medizinische oder juristische Beratung. Sprachmodelle halluzinieren — auch fein-getunte „Medizin“-Modelle. Für regulierte Einsätze gelten Validierung, menschliche Aufsicht und Recht (z. B. Medizinprodukte/MDR, DSGVO). Benchmarks wie MedQA belegen Testleistung, nicht klinische Eignung.

Important (medical/legal): Nothing on this page is medical or legal advice. Language models hallucinate — fine-tuned "medical" models included. Regulated use requires validation, human oversight and compliance (e.g. medical-device/MDR, GDPR). Benchmarks like MedQA show test performance, not clinical fitness.

Block BBlock BMomentaufnahmeToday snapshot

Diese Tabellen kommen aus data/models-snapshot.json — erzeugt von einem Skript, nicht live im Browser (siehe Block C). Zahlen sind eine Momentaufnahme mit Quelle und Stand, keine ewige Wahrheit.

These tables come from data/models-snapshot.json — produced by a script, not fetched live in the browser (see Block C). Numbers are a snapshot with source and date, not eternal truth.

Lade Momentaufnahme …

Loading snapshot …

Block CBlock CQuellen-RegisterSource register

Wo die Zahlen herkommen — und wo jede Quelle „lügt“, also systematisch in die Irre führen kann. Das Skript zieht nur öffentliche Endpunkte; es baut die Ranglisten nicht nach.

Where the numbers come from — and where each source "lies", i.e. can systematically mislead. The script pulls public endpoints only; it does not rebuild the leaderboards.

OpenRouter — /api/v1/models

Katalog mit Preis und Kontextlänge über viele Anbieter. Wir nutzen ihn für Modell-Liste, Preise und Kontext.

Catalogue with price and context length across many providers. We use it for the model list, prices and context.

⚠ Lügt bei: Qualität — es sagt nichts darüber, wie gut ein Modell ist; Preise sind Router-Preise und können vom Anbieter abweichen.

⚠ Lies about: quality — it says nothing about how good a model is; prices are router prices and can differ from the provider.

Hugging Face Hub — /api/models

Öffentliche Meta-Daten offener Modelle: Downloads, Likes. Wir nutzen sie als Adoptions-Signal (offene Gewichte, Medizin).

Public metadata for open models: downloads, likes. We use it as an adoption signal (open weights, medical).

⚠ Lügt bei: Downloads ≠ Qualität. Kleine Modelle (0,6B) und Quantisierungen führen die Liste an, weil sie oft geladen werden — nicht weil sie besser sind.

⚠ Lies about: downloads ≠ quality. Tiny models (0.6B) and quantisations top the list because they're pulled often — not because they're better.

LMArena / Arena.ai Elo

Menschliche Blind-Votes → Elo. Keine offizielle API; wir spiegeln optional einen Community-Mirror. Fällt der aus, bleibt die Spalte leer statt zu raten.

Human blind votes → Elo. No official API; we optionally mirror a community copy. If it fails, the column stays empty instead of guessing.

⚠ Lügt bei: Stil schlägt Substanz. Elo belohnt angenehme Antworten; Prompt-Mix und Sprache verzerren.

⚠ Lies about: style beats substance. Elo rewards pleasant answers; prompt mix and language skew it.

Artificial Analysis — artificialanalysis.ai

Aggregierter „Intelligence Index“ plus Preis/Geschwindigkeit. Guter Kompass zum Nachlesen.

Aggregated "Intelligence Index" plus price/speed. A good compass to read up on.

⚠ Lügt bei: der Index ist ein gewichteter Mix — die Gewichte sind eine Meinung, nicht dein Use Case.

⚠ Lies about: the index is a weighted blend — the weights are an opinion, not your use case.

SWE-bench / LiveCodeBench / LiveBench

Coding- und kontaminations-arme Benchmarks. Für die Coding-Klasse relevant.

Coding and contamination-resistant benchmarks. Relevant for the coding class.

⚠ Lügt bei: Harness und Prompt-Gerüst beeinflussen SWE-bench stark — „% gelöst“ hängt am Setup, nicht nur am Modell.

⚠ Lies about: harness and scaffolding heavily affect SWE-bench — "% solved" depends on the setup, not just the model.

Open Medical LLM Leaderboard + Medmarks

MedQA, PubMedQA & Co. für die Domäne Medizin. Nur damit lässt sich „nachweislich besser“ belegen.

MedQA, PubMedQA & co. for the medical domain. This is what "provably better" needs.

⚠ Lügt bei: Multiple-Choice-Testleistung ist keine klinische Sicherheit. Gut im Test ≠ gut am Patienten.

⚠ Lies about: multiple-choice test scores are not clinical safety. Good on the test ≠ good with a patient.

Mehr lesen: Read more: Stanford HELM Scale SEAL Vals.ai LLM Stats

Unsicher, welche Klasse zu deinem Projekt passt? Kurze Mail genügt.Unsure which class fits your project? A short mail is enough.