Block ABlock ADer LeitfadenThe guide
Generische Probleme
Generic problems
Nimm ein aktuelles Frontier-ModellTake a current frontier model
Für die allermeisten Aufgaben: ein aktuelles Frontier-Modell eines großen Labors (OpenAI, Anthropic, Google, xAI). Lauf nicht jede Woche dem Platz-1-Modell hinterher. Stabilität, gutes Tool-Use, langer Kontext und ein guter Agent drumherum schlagen zwei Punkte auf einem Leaderboard.
For the vast majority of tasks: a current frontier model from a big lab (OpenAI, Anthropic, Google, xAI). Don't chase the rank-1 model every week. Stability, solid tool-use, long context and a good agent around it beat two points on a leaderboard.
Sehr spezifische Domänen
Very specific domains
Frontier zuerst — Spezialmodell nur mit BelegFrontier first — specialist only with proof
Für enge Domänen (Medizin, Bio, Recht) startest du trotzdem mit einem Frontier-Modell als Default. Dann prüfst du, ob ein dediziertes oder fein-getuntes Modell nachweislich besser ist — über MedQA, PubMedQA, Medmarks oder deine eigenen Testdaten. Nimm das Spezialmodell nur, wenn Benchmark und dein Use Case es stützen — nicht, weil „Med“ im Namen steht.
For narrow domains (medicine, bio, law) you still start with a frontier model as the default. Then you check whether a dedicated or fine-tuned model is provably better — via MedQA, PubMedQA, Medmarks or your own test data. Pick the specialist only if the benchmark and your use case support it — not because the name contains "Med".
Immer Quelle und Datum mitzeigen. Ein Benchmark-Wert ohne „woher“ und „wann“ ist wertlos. Und: Leaderboards sind ein Filter, keine Wahrheit — sie messen enge Aufgaben unter idealen Bedingungen, nicht deinen Alltag.
Always show source and date. A benchmark number without "where from" and "when" is worthless. And: leaderboards are a filter, not truth — they measure narrow tasks under ideal conditions, not your day-to-day.
Die ModellklassenThe model classes
Frontier (proprietär)Frontier (closed)
GPT · Claude · Gemini · Grok
Default für die meisten Aufgaben.
Default for most work.
Frontier (offene Gewichte)Frontier (open-weight)
Llama · Qwen · DeepSeek · Mistral · GLM · Kimi
Self-Hosting, Datenkontrolle, Anpassung.
Self-hosting, data control, customisation.
Coding / AgentenCoding / agents
Codex · Qwen Coder · Devstral · Codestral
Code schreiben, Tools nutzen, Agenten.
Writing code, using tools, agents.
Reasoning / STEMReasoning / STEM
o-Serie · R1 · „Thinking“-Varianten
Mehrschritt-Logik, Mathe, Beweise.
Multi-step logic, math, proofs.
Günstig / schnellCheap / fast
mini · flash · haiku · nano · lite
Hohes Volumen, niedrige Latenz, wenig Kosten.
High volume, low latency, low cost.
Domäne: MedizinDomain: medical
Bio-Medical-Llama · II-Medical · …
Fein-getunt — nur mit Beleg statt als Default.
Fine-tuned — only with proof, not as a default.
Bild/Video (Imagen, Veo, Sora, Lyria) sind eine eigene Welt und hier bewusst ausgeklammert.Image/video (Imagen, Veo, Sora, Lyria) are their own world and deliberately left out here.
Wichtig (Medizin/Recht): Nichts auf dieser Seite ist medizinische oder juristische Beratung. Sprachmodelle halluzinieren — auch fein-getunte „Medizin“-Modelle. Für regulierte Einsätze gelten Validierung, menschliche Aufsicht und Recht (z. B. Medizinprodukte/MDR, DSGVO). Benchmarks wie MedQA belegen Testleistung, nicht klinische Eignung.
Important (medical/legal): Nothing on this page is medical or legal advice. Language models hallucinate — fine-tuned "medical" models included. Regulated use requires validation, human oversight and compliance (e.g. medical-device/MDR, GDPR). Benchmarks like MedQA show test performance, not clinical fitness.