Files
roboco/docs/rag/workflows/kb-search.md
T
4fb0059556 fix: journals/learnings never reached the RAG corpus + git-readonly slug 404s (#339)
* fix(rag): per-index chunk floors — journals and learnings were never indexed

The global 200-char garbage floor (sized for code/doc chunks) discarded
every templated journal note and most distilled org-memory lessons,
silently: ingest returned success with zero chunks, so agent journals
and learnings were never retrievable via RAG. IndexConfig now carries a
per-type min_chunk_length (journals 40, learnings 80, others unchanged).

* fix(mcp): git-readonly tools default project_slug from the container env

Agents 404ed /api/git/status with 'Project not found: roboco' — the
tools made the LLM supply the slug and six doc examples taught a slug
that matches no registered project. The tools now fall back to the
ROBOCO_PROJECT_SLUG the orchestrator already injects, and the stale
examples are corrected.

* feat(rag): startup backfill re-ingests zero-chunk journals and learnings

Before the per-index chunk-floor fix, ingest() returned success with
chunk_count=0 for undersized content: every historical journal entry and
distilled learning below the (then-global) 200-char floor was durably
recorded in journal_entries but silently never got a chunks_journals /
chunks_learnings row, and no exception meant the existing dead-letter
(rag_index_failures) never saw it either.

Extends the startup reconcile (roboco/api/app.py _reconcile_rag_indexes)
with a new pass: backfill_unindexed_journals (roboco/services/
rag_index_failures.py) queries journal_entries for rows missing from each
vector table and re-ingests them through the same live code paths
(_reindex_journal_entry / record_learning).

Journals and learnings are backfilled independently since a LEARNING entry
can clear the (lower) JOURNALS floor while still failing the (higher)
LEARNINGS floor — a learning's doc_source is a content hash, not the entry
id, so presence there is checked by hashing each candidate the same way
LearningsIndexPlugin.record_learning does and batch-querying chunks_learnings
for those exact sources.

Bounded to 200 rows per pass per boot (converges over restarts on a larger
backlog) and best-effort per row (one failure never aborts the pass). Rows
still under the current floor are excluded by a length filter in the SELECT
so they are never retried forever, and private entries are excluded from
the JOURNALS pass exactly like the live indexing path.

* test(rag): scope backfill assertions to their own rows

---------

Co-authored-by: Renn F <rennf93@users.noreply.github.com>
2026-07-08 16:03:00 +02:00

103 lines
2.4 KiB
Markdown

# Knowledge Base Search
**ALL agents have access to KB/RAG tools.** These are automatically available.
## Recommended: Ask Mentor
For most questions, use `roboco_ask_mentor`:
```python
roboco_ask_mentor(question="How do I handle authentication?")
```
It searches ALL knowledge sources and supports follow-up questions.
## Search Types
| Tool | Purpose | Best For |
|------|---------|----------|
| `roboco_ask_mentor` | Conversational help | **Most questions** |
| `roboco_kb_search` | Semantic search | Browsing, exploration |
| `roboco_rag_query` | AI-synthesized answer | Quick answers |
## Semantic Search
```python
roboco_kb_search(
query="rate limiting redis implementation",
top_k=5, # Results to return
project="roboco-api", # Optional project filter
index_types=["code", "docs"] # Filter by type
)
```
Returns similar content - not just keyword matches.
## RAG Query (AI Answer)
```python
roboco_rag_query(
query="How does authentication work in this codebase?",
top_k=5
)
```
Returns AI-synthesized answer with citations.
Good for:
- "How does X work?"
- "What pattern should I use?"
- "What decisions were made about Y?"
## Mentor (Conversational)
```python
# First question
response = roboco_ask_mentor(
question="How do I handle authentication?",
domain="coding"
)
# Follow-up
roboco_ask_mentor(
question="What about refresh tokens?",
conversation_id=response["conversation_id"]
)
```
## Index Types
| Type | Content |
|------|---------|
| `code` | Source files |
| `docs` | Documentation |
| `journals` | Agent journal entries |
| `errors` | Error patterns & fixes |
| `standards` | Coding rules |
| `decisions` | Architectural decisions |
| `reviews` | Code review patterns |
| `learnings` | Captured learnings |
## Before Starting a Task
Always search first:
```python
roboco_kb_search(query="implementing rate limiter")
# Journal entries are part of the KB — filter to them with index_types:
roboco_kb_search(query="rate limit decisions", index_types=["journals", "decisions"])
```
This helps you:
- Avoid repeating mistakes
- Find proven patterns
- Learn from others' experiences
## Proactive Context
System auto-provides context when you claim:
```python
roboco_get_proactive_context(task_id)
# Returns: similar_tasks, relevant_learnings, code_patterns,
# applicable_standards, recent_decisions, known_issues
```