Appearance
The same wall, over and over
Every time you solve a real problem with an AI assistant, the solution has a half-life of about one session. Question, answer, problem solved, and the conversation sinks into chat history. The next person hits the same wall and runs the same loop: same investigation, same fix, same tokens burned twice.
That waste is the premise behind Shared Knowledge, an open-source MCP server built for a dev.to weekend challenge. Its tagline is the whole argument: "Don't just get the answer. Give the answer back." It turns a solved problem into a Markdown article, opens a pull request, and publishes only after a human approves. But the project is also part of a bigger shift. Over the last 18 months, the agent stack has crystallized into three portable layers: skills, MCP servers, and the harness. OpenAI has spent the last 17 months unbundling Codex layer by layer, and the pieces are standard enough that they're reshaping how teams build agentic systems.
Three layers, one stack
The model is only half of an agent. The other half is everything around it: the loop keeping it pointed at a goal, the context management stopping it from drowning, the tools it can call, the subagents it spins up. DeepSeek's engineers compressed that into an equation when they shipped DeepSeek Harness: Agent = Model + Harness.
The equation is showing up in practice. The harness is the layer OpenAI peeled out of Codex and now sells as a hosted API. The skill is the layer that book-to-skill, Claude-Red, and a dozen other repos are standardizing around one file format. MCP is the glue connecting all of it to whichever client you already use.
The model decides what the agent should attempt. The harness decides whether it finishes.
Skills: behavior, packaged
A skill is a SKILL.md file: plain Markdown that primes a model with methodology, decision rules, and structure, loaded only when relevant. The format has become the closest thing the agent world has to a plugin standard. GitHub Copilot CLI, Amp, Claude Code, Codex, Cursor, and Hermes Agent all read the same file. Install once, use everywhere.
The range of things people are packaging is wider than you'd expect. i-have-adhd encodes output conventions, not knowledge: ten rules demanding you lead with the next action, number multi-step tasks, give time estimates in minutes, cap lists at five items, and skip the closers. The first answer it produced for me started with the concrete next step instead of a recap, and that's the whole pitch. no-ai-slop attacks the opposite failure, the writing that AI smooths until it sounds like a corporate press release. It checks for 20+ patterns, including binary contrasts, throat-clearing openers, and fake-profound endings, then lists what it changed instead of silently rewriting. A second mode quotes every pattern it finds without making claims about who wrote the text, and a third generates the cringiest AI slop it can on demand, as satire.
Claude-Red packages offensive security methodology: roughly 80 skills across 20+ categories, from SQL injection to EDR evasion to ADCS abuse. Triggers load skills on demand, so a conversation about wireless testing pulls in WPA2 and BLE methodology without paying context for the rest. google-gemini/gemini-skills is the vendor-curated case. Google's evaluation found that adding the gemini-api-dev skill pushed correct API code generation to 87% with Gemini 3 Flash and 96% with Gemini 3.1 Pro. The reason is blunt: models don't know about themselves at training time, and SDKs change faster than retraining cycles.
| Skill | What it packages | Key detail |
|---|---|---|
| book-to-skill | Books and long docs, distilled into structured skills | 24x-51x fewer tokens per query |
| i-have-adhd | Output format rules for coding assistants | Lead with the next action, cap lists at 5 |
| no-ai-slop | Writing quality checks | Flags 20+ slop patterns, logs its edits |
| Claude-Red | Offensive security methodology | ~80 skills, trigger-loaded |
| gemini-skills | Gemini API and Live API knowledge | 87% Flash / 96% Pro correct API code |
Quick Take: the same skill file now moves between Claude, Codex, Cursor, and Gemini unchanged, which makes skills less like prompts and more like portable plugins.
The token economics of skills
book-to-skill makes the strongest quantitative case for the format. Point it at a PDF or a docs folder and it distills the content into a structured skill: a core SKILL.md of roughly 4,000 tokens, one chapter file per chapter at about 1,000 tokens each, plus a glossary, a patterns file, and a cheatsheet. Chapters load on demand. Ask about chapter 7 and the agent reads chapter 7, not the whole book.
The measured result is 24x to 51x fewer tokens than dumping the book into context to answer a single question. The mechanics explain why. An agent reading a raw PDF doesn't just read it; it re-fetches the table of contents, backtracks, and re-processes everything on every turn. book-to-skill pays that structuring cost once, at conversion, so queries stay proportional to the answer.
That changes what "loading context" means. A 200-page technical book would swallow most context windows. A distilled skill is a few thousand tokens sitting on disk until needed. For a team that re-reads the same runbooks, architecture decision records, or compliance docs, this is the difference between asking a question and re-reading the manual. If you reopen a document often enough to wish you'd memorized it, it's a candidate.
MCP: plumbing for shared knowledge
MCP is the layer that lets agents reach outside the chat window. The clearest demonstration I've seen is Shared Knowledge, an MCP server that turns a solved problem into a reusable article. The flow is deliberately minimal:
The project made a few decisions that look right for an MVP. GitHub is the single source of truth: the knowledge base is plain Markdown, so git already handles versioning, history, branches, diffs, and review. Search is field-weighted keyword search instead of a vector database, because at four articles an embedding pipeline would add weight, not intelligence. And the hardest boundary is enforced: nothing publishes without human review. The conversation stays private. Only the extracted solution is offered up. Only the merged PR becomes public.
The first published contribution shows the loop working end to end. A pyproject.toml optional dependency that crashed the import chain became an article, went through an automated Copilot review that caught YAML formatting issues, got fixed, merged, and ended up as both a docs page and an ElevenLabs audio file.
The architectural call I respect most is what the author gave up. The first version ran Gemini inside the server to structure the article. That worked, but it turned the server into a monolith tied to one vendor, which defeats the point of MCP. The author ripped Gemini out mid-build, exposed a knowledge_article_guidelines prompt instead, and let the calling assistant structure its own output. The server only validates and publishes. The explicit cost was giving up eligibility for the Best Use of Google AI prize category. Interoperability won because it's the point of the protocol.
One detail caught my attention: the second contribution went through GitHub Copilot Chat in Agent mode, and the assistant spontaneously queried the existing knowledge base to check for duplicates before publishing. Nothing hard-coded that behavior. The server provides tools; the model decides how to use them. That's the MCP philosophy in a single observation.
What the community is saying: the Stack Overflow comparison keeps coming up. When I read the comments on the project, one framing stuck: a searchable base of working setups is a template-grab system for project-scoped problems, like how to wire a CLAUDE.md for a Blazor project. The difference from a Q&A thread is that the answer arrives already structured and reviewed, not buried in twelve replies of varying quality. The recurring fear, which I share, is that a user-fed base only stays useful if people actually give answers back.
Scrapling shows MCP serving a different purpose: an MCP server that lets agents scrape through one-shot or session-based tools, with pages narrowed by CSS selectors and stripped of prompt-injection content before the model sees them. That last detail matters. When you hand an agent a scraper, you're also handing the open web a vector for prompt injection. Sanitizing pages pre-read is the difference between a useful tool and a liability.
The harness: Agent = Model + Harness
The harness is the runtime layer. It handles long-session context compression, tool scheduling, subagent coordination, authentication, and state. It's the part that decides whether an agent actually finishes a task instead of producing a confident paragraph about it.
OpenAI has spent a year peeling this layer out of Codex and selling it separately. The timeline is the clearest product arc in the space:
| Date | Release | What changed |
|---|---|---|
| April 2025 | Codex CLI open-sourced | Harness logic public, self-hosted |
| May 2025 | Codex cloud | Task spawns an isolated sandbox, parallel tasks |
| October 2025 | Codex SDK | Drive the agent from TypeScript, structured output, pause/resume |
| February 2026 | Codex App Server | JSON-RPC interface over the full harness |
| August 2026 | Open Codex Harness | CLI, SDK, App Server unified as a platform |
| September 2026 | Agents API public beta | Task, model, tools, runtime; OpenAI hosts the harness |
The September step is the important one. You tell the API four things: the task, the model, the tools, and the runtime. OpenAI handles context management and subagent coordination. There's no separate platform fee; you pay for model tokens and tools, and compute if you use OpenAI's sandbox. You can also point it at your own infrastructure, or Cloudflare, E2B, or Modal. The harness became a hosted service: SaaS, but for the agent loop instead of software.
Anthropic got there first in most respects. Claude Agent SDK shipped in September 2025, exposing the tooling, context management, permissions, and subagent capabilities behind Claude Code. Claude Managed Agents arrived in April 2026, splitting session, harness, and sandbox into independent layers, with Anthropic hosting the harness and long-running tasks while the sandbox stays swappable. Google joined at I/O 2026 with Gemini API Managed Agents running the Antigravity harness as a managed service. Sequencing matters less than direction: every major model company now sells the harness.
Key numbers: the Codex unbundling took about 17 months from CLI to Agents API. The Agents API adds no platform fee; billing runs on model tokens, tools, and any sandbox compute you consume. Across the skills layer, book-to-skill measures 24x-51x token savings per query, and Gemini's API skills land at 87% (Flash) and 96% (Pro) correct API code.
Two philosophies, one prize
Two philosophies are forming at the harness layer, and a 36Kr analysis framed it as a fork: DeepSeek is building the Linux of harnesses, OpenAI is building the AWS.
DeepSeek Harness is a modular open framework where everything is a plugin: model, tools, skills, session, sandbox, storage, agent loop, scheduling, even the UI. Assemble your own, run it anywhere, swap every part. OpenAI's Agents API is the opposite bet: hosted, maintained, continuously upgraded, with the wiring hidden. Give them the task, the budget, and the requirements; harness maintenance becomes someone else's problem.
Both are chasing the same prize: becoming the default harness developers reach for. That prize is bigger than any single model win, because the model sets the ceiling on what an agent can do while the harness decides whether the work gets done.
And the contest is shifting from intelligence to execution. An agent that actually does work needs access to email, documents, meetings, accounts, and permissions. The platform companies that accumulated those assets over two decades are sitting on a keychain of APIs and data, and model companies have to connect to the world one integration at a time. The office-agent wars playing out in China are a preview: the winners may be the companies that already own the enterprise data and tooling, not the ones with the best benchmark scores.
Google is the extreme full-stack case. It owns TPUs, cloud, Gemini, Search, Workspace, Chrome, and Android, which together form a complete environment an agent can live in. Google has also started consolidating its agent infrastructure behind the Antigravity harness. But the user-facing story is still scattered across Gemini Spark, Workspace Studio, Antigravity, Gemini Enterprise, and the agents in Search. Google has all the pieces and hasn't yet shipped the one simple answer. If it does, the competitive picture shifts.
What this means for builders
If you're shipping agentic products, stop building the harness from scratch. The managed options exist and they're good enough that your scarce engineering time belongs in the layers that differentiate your product: the skills that encode your domain knowledge, and the MCP servers that connect your tools. A weekend project like Shared Knowledge is proof that a single developer can now build a useful agent-adjacent service without touching the loop underneath.
If you're building internal knowledge systems, the cross-host skill standard is your distribution channel. One install covers Copilot CLI, Amp, Claude Code, Codex, and Cursor. The conversion workflow from book-to-skill applies to anything you re-read: runbooks, architecture decision records, onboarding guides, brand guidelines, RFCs. Fold a whole docs folder into one skill and ask questions while you code.
If you're picking between open and hosted harnesses, frame it as an operational decision. Open harnesses give you control and hand you the maintenance. Hosted ones move that burden to the vendor and tie you to their roadmap. Neither choice is wrong, but it's a strategy decision, not a technical one.
Common Pitfalls
Treating a skill as a context dump. Stuffing a 60-page document into one SKILL.md burns tokens on every turn. The structure is the point: a small core file, chapter files loaded on demand, a cheatsheet for quick reference. Otherwise you've just renamed the PDF. Structure, not summary.
Putting a model inside your MCP server. The first version of Shared Knowledge did exactly this, and the author tore it out mid-build. A server should expose tools and prompts; the calling assistant brings the model. The moment you hard-code a vendor model into the server, you've built a monolith, not a protocol.
Publishing without the human gate. The rule that keeps shared knowledge trustworthy is that nothing goes public without review. Four reviewed articles in the first weekend is exactly the right pace for a system whose whole value is trust. The moment auto-publish enters the pipeline, the base fills with unvalidated garbage.
Sending Markdown straight to a TTS engine. Headings like ## Problem run directly into the following text with no prosodic break, and the audio comes out flat. Markdown is built for Git and diffs, not for speech. The fix is an intermediate narration-script step: parse the structure, insert pauses and transition cues, then synthesize.
Trusting generated outputs without verification. A code review tool flagged a convincing authentication bug, complete with an excerpt and line number, on a line that didn't exist in the repo. A deleted MP3 with no manifest update had CI concluding the audio was current. And an empty ELEVENLABS_MODEL_ID in CI silently overrode the code's default, because os.environ.get() only falls back when the key is missing, not when it's set to an empty