Skip to content

Retrieval Beats Generation: What Four RAG and In-Context Approaches Agree On

#retrieval-augmented-generation #in-context-learning #ontology-learning #machine-translation #memory-architecture

LLMs sound certain about everything. That certainty evaporates the moment they need to ground output in a specific schema, a specific document, or a specific project's history. The usual fixes all orbit one idea: change what goes into the context window.

Four recent sources attack this from different angles. An ontology learning pipeline for the LLMs4OL challenge retrieves demonstration triples to keep outputs format-consistent. A machine translation framework extracts parallel fragments from exemplars and uses them as reasoning traces. A batch prompting paper separates reasoning from symbol grounding to make inference efficient without losing accuracy. And a developer on dev.to argues that memory should live outside the model entirely, as a shared external layer with trust states.

What's interesting isn't that retrieval helps. It's how differently each approach decides what to retrieve. Documents, fragments, batch assignments, project history. And all of them hit the same wall: the hard part is choosing the right thing to put in context, not generating the answer.

Quick Take: RAG, in-context reasoning, batch prompting, and external memory all solve the same core problem: deciding what deserves to be in the context window. The hard part is always retrieval, not generation.

Retrieval-Augmented Few-Shot for Ontology Learning ​

Ontology learning turns raw text into structured triples: entities, their types, and the relations between them. LLMs are decent at this until they aren't. They hallucinate domain terms, produce inconsistent formats, and default to hierarchical relations over associative ones. The LLMs4OL 2026 Challenge asked teams to handle two jobs: the End-to-End Flagship Task (Task A) and the Ontology Extension Reuse Task (Task B).

The system here uses Qwen2.5-14B-Instruct as the generator and all-MiniLM-L6-v2 as a retriever. It pulls the top-5 demonstrations for Task A and the top-2 for Task B. A left-truncated context window preserves the task instructions even when the prompt grows. For Task B, generated triples pass through a deterministic vocabulary filter that keeps only triples where at least one endpoint belongs to the sample's closed term/type vocabulary, and removes duplicates against the initial ontology.

Results are strong on what the vocabulary covers. Task B hits 0.8692 Semantic Graph Similarity, 0.9200 Term-Typing F1, and 0.8540 Taxonomy Discovery F1. Task A gets 0.7416 SGS.

But no non-taxonomic relations were extracted. The relation vocabulary is taxonomy-oriented, and the retriever never surfaced an associative relation example. The model has no evidence of what an associative triple looks like, so it never produces one. That's a closed-vocabulary ceiling, not a generation failure.

Key numbers from the LLMs4OL system:

  • 14B parameters is small enough to run on a single high-end GPU with quantization.
  • top-5 vs top-2 demonstrations: Task A needs more evidence, Task B gets away with a couple of examples.
  • 0.9200 Term-Typing F1 means the type/term classification subtask is nearly solved.
  • zero non-taxonomic relations extracted is the exact failure mode you'd predict from a taxonomy-only vocabulary.

Fragments as Reasoning Traces for Translation ​

Machine translation with LLMs has a similar problem. In-context examples help, but the model often copies surface style or misses the correspondences that actually matter. The paper proposes a fragment-based reasoning framework. Instead of handing the model full exemplar pairs and hoping it extracts the right alignments, the model first extracts parallel source-target fragments from retrieved similar exemplars. Those fragments become intermediate reasoning traces. Only then does it generate the final translation.

Training relies on distillation. A large teacher model produces silver fragments and drafts, which train the student Qwen3 model. The experiments cover 6 languages and up to 5 domains per language. Fragment-based MT beats standard k-shot and basic drafting across the board.

The abstract doesn't publish exact margins. What matters is that the win holds across multiple languages and domains. This isn't a trick that works on one test set. The trick is forcing the model to articulate which pieces of the exemplar actually transfer before it writes anything.

Cascaded Batch Prompting ​

Batch prompting makes inference cheaper by stuffing multiple instances into one prompt. The catch is that downstream performance becomes unpredictable. The proposed solution, cascaded batch prompting, splits the work into two stages. In the first stage, the model reasons about the batch of questions without worrying about aligning its reasoning to specific symbols. In the second stage, it grounds that reasoning to the actual answer choices or labels. Complex reasoning and symbol grounding get disentangled.

The method beats the standard single prompting baseline while achieving a speedup proportional to batch size. It claims a new point on the accuracy-throughput Pareto frontier. For production systems that balk at the latency of one-inference-per-sample pipelines, this is the most immediately useful result of the four.

Notice the pattern. The failure mode it fixes is not the model's reasoning ability. It's the model losing track of which symbol belongs to which case when everything runs in one pass. That's a retrieval problem in disguise: the model needs to retrieve the right assignment between internal reasoning and external labels.

Memory Moves Out of the Model ​

I found the essay on dev.to by Marcos Somma after reading the papers, and it reframes the whole cluster. He argues that AI memory shouldn't belong to the model at all. It should belong to the system around the model. His prototype is a shared external memory layer: plain Markdown files with structured metadata, split between general and project-specific knowledge. Access goes through MCP, and each session sees a small index first. Full content enters context only when the agent explicitly asks for it.

The key move is the boundary. The model isn't the memory store, the client isn't, and the MCP server isn't. Memory exists independently, so Claude can write something today and GPT can read it tomorrow. A local model can challenge it next week.

Every memory carries a trust state. An unreviewed memory is a report, not an instruction: "Agent X said this was true at time Y." If an AI edits a reviewed memory, the review badge disappears. Review is a promotion mechanism, not a write path. Agents create low-trust memories cheaply; only stable project guidance consumes human review. Otherwise you've just rebuilt the documentation bottleneck one layer later.

He also refuses to auto-delete old memories. Age is not truth. A four-year-old architectural decision can still explain why half the system looks the way it does. A memory created this morning can already be nonsense. So the hub flags stale memories and explicitly tells the agent to verify claims against current code. A cache asks whether a value can be reused. Memory asks how much you should believe it now.

The deeper argument is economic. He calls it the exploration tax. Every fresh session re-discovers what the team already learned: we already tried that migration, we cannot change that response format, there's a reason this interface looks wrong. Token usage may stay flat or even go up because the agent reads memories on top of source files. But human corrections and repeated architectural mistakes drop. Engineer-hours cost more than extra tokens.

What Connects Them: The Retrieval Economy ​

Every one of these approaches changes how context gets selected. The ontology pipeline selects demonstrations. Fragment-based MT selects aligned pieces of exemplars. Batch prompting selects how to bind reasoning to symbols. Memory selects which historical facts matter. In every case, you're building a retrieval economy where the index is cheap and the details cost context.

The cost of retrieving the wrong thing is visible in all four. The ontology system's zero non-taxonomic relations are direct evidence: the retriever never returned an associative relation example, so the model never generated one. Memory has the same problem. A memory that exists but is never retrieved is functionally forgotten. Pulling irrelevant memories is just context pollution with extra steps.

ApproachWhat it retrievesProblem it solvesFailure mode
RAG for ontology learningTop-k demonstration triplesFormat consistency, hallucinated termsClosed vocabulary misses relation types
Fragment-based MTParallel source-target fragmentsTranslation alignmentGood retriever required to find useful exemplars
Cascaded batch promptingBatch-to-symbol assignmentsInference inefficiencyGrounding errors when reasoning and symbols merge
External memoryHistorical project decisionsRediscovering decisions across sessionsStale or unproven memories mislead agents

All four treat retrieval precision and recall as core evaluation criteria, not side metrics. The ontology paper explicitly points to the missing relation type as a retrieval failure. The memory essay says retrieval precision and recall have to be part of the evaluation, not something smuggled into the phrase "the agent decides."

Common Pitfalls ​

Don't make these mistakes.

  • Don't assume a bigger context window solves grounding. The ontology system's left-truncation trick exists because long prompts bury the instructions. A larger window just gives you more space to bury the signal.
  • If you have a closed vocabulary, retrieval will never surface what the vocabulary doesn't contain. When your relation set excludes associative relations, no number of few-shot examples will produce them. Fix the vocabulary first.
  • Fragment-based retrieval doesn't tolerate copy-paste. If you extract fragments naively, you get noise. The MT paper distills fragments from a large teacher model. Skip that step and you'll wonder why the reasoning traces are garbage.
  • Batch prompting with very different instance lengths unbalances attention. Cascaded batch prompting works because it separates reasoning from grounding. If you merge the stages back together, you reintroduce exactly the unpredictability you were trying to kill.
  • For memory layers, never treat an old memory as automatically true. If the agent can't verify a memory against code or plans, it can import yesterday's confident wrong conclusion into today's session. Provenance and stale flags are not optional.

One thing to remember: retrieval precision is the product. Every one of these techniques delivers its value at the retrieval step, not the generation step. If your retriever surfaces the wrong exemplar, the wrong fragment, or the wrong memory, no amount of prompt engineering will fix it.

The Bottom Line ​

If you're building a system that needs consistent schema-level output like ontologies or knowledge graphs, adopt retrieval-augmented few-shot with a retriever that can actually find examples of the relation types you care about. A closed vocabulary will choke every other part of the pipeline.

If you're doing machine translation in production, the fragment-based reasoning approach is worth the engineering cost, but only if you're willing to build the distillation pipeline. Standard k-shot is easier to demo and won't give you the same alignment gains.

If you're designing an agent platform with persistent memory, skip the model-owned memory and put it in a shared layer with provenance. Watch the exploration tax: measure human corrections and repeated wrong paths, not just token usage. Token savings may be neutral while the real win is stopping settled decisions from being relitigated.