Research & Insights

We publish what we learn. Benchmarks, frameworks, and principles from building AI systems for real companies.

Applied AI

AI Reaction Conditions Extraction: Model Benchmarks from a BioTech Backfill

We benchmarked four AI approaches against 1,364 chemicals worth of handwritten scans, patents, and protocols. The cheapest model failed completely. The most expensive tied with the second-cheapest. The real bottleneck was a 300x S3 slowdown nobody expected.

Read
Strategy

Data Sovereignty as an Operating Model

Sovereignty is not a compliance checkbox. Organizations that treat it as an architecture discipline accumulate control, auditability, and compounding institutional knowledge. Organizations that rent their intelligence layer leak all three.

Read
Agent Skills

SkillOpt in Production: Optimizing Agent Skills Beyond Benchmarks

We applied gradient descent to markdown skill files. Three production skills, +18.4% average improvement, 100% cross-model portability, convergence in 2 steps. First practitioner validation of the SkillOpt framework.

Read
Inference

We Tested Gemma 4 for Production Agent Work. Here's What Held Up.

32 tests across text and vision. 100% pass rate. Google's 26B MoE model benchmarked for the tasks AI agents actually do: classification, structured output, code generation, chart reading, and anomaly detection.

Read
Inference

Nemotron 3 Nano: The Small Model That Doesn't Act Small

NVIDIA's 30B/3B-active hybrid Mamba-Transformer-MoE delivers 3.3x throughput over comparable models. We break down the architecture, benchmark it against Qwen3.5, and map out when to use it versus when not to.

Read
Memory

We Benchmarked 5 AI Agent Memory Systems

50-question eval framework, 5 memory solutions tested, real production data. Cognee + hook-based architecture delivered 13x improvement over baseline.

Read
Inference

vLLM 0.17: The Flag That Changed Everything

We upgraded vLLM and performance got worse. Then we added one flag and latency dropped 6x. Default v0.17 with --enforce-eager is slower than v0.16. Here is why, and exactly what to do about it.

Read
Inference

GPT-oss 120B vs Qwen3-30B: Production Benchmark

OpenAI released their open-weight flagship. We ran it against our production model on real business tasks. The 30B model won on quality, matched on speed, and costs less than half to run.

Read
Inference

Qwen3.5 vs Qwen3: Day-One Benchmarks

We ran our full production benchmark suite against Qwen3.5-35B-A3B the day it released. Quality matched at 11/11. Throughput was half. Here is why that gap exists and exactly when to switch.

Read
Inference

State of Self-Hosted Inference: February 2026

We run our own inference infrastructure for privacy and control. Here is what the self-hosted model landscape looks like right now, what we run in production, and the real benchmark numbers behind our decisions.

Read
Strategy

The AI-Native Framework

What does AI-native actually look like? We mapped 23+ engagements across 9 industries and pulled out the patterns. Six building blocks, four transformation steps, and what each vertical looks like when it gets it right.

Read
Data

Inference Benchmarks

We benchmark AI models on real business tasks: classification, email drafting, code generation, summarization, and more. Comparing BF16, GPTQ-Int4, and Instruct-tuned variants on dedicated A100 hardware.

Read
Design

Product Design Principles

How we think about product design. Principles from designers at Cursor, Notion, and Stripe applied to AI product development. The four questions every user asks, CTA hierarchy, and the anti-pattern checklist.

Read

We use cookies to improve your experience. Cookie Policy · Privacy Policy