AxilJS RAG Pipeline for Retrieval-Augmented Generation
Build retrieval-augmented generation (RAG) applications with AxilJS using document indexing, vector search, AI-generated answers, source retrieval, and streaming.
AxilJS RAG Pipeline
AxilJS provides a Retrieval-Augmented Generation (RAG) pipeline for building AI applications that retrieve relevant information from indexed documents and use that information to generate contextual answers.
The RAGPipeline combines an AI provider with a VectorStore, allowing applications to index text, perform retrieval-based queries, return source documents, and stream RAG responses.
What Is RAG?
Retrieval-Augmented Generation combines two stages:
- Retrieval — Find relevant information from indexed documents using vector search.
- Generation — Pass the retrieved information to an AI provider to generate an answer.
A typical AxilJS RAG workflow looks like this:
This approach allows an AI application to answer questions using information retrieved from a specific document collection.
Setup
Import RAGPipeline, VectorStore, and loadFromText from @axiljs/ai.
The provider is the configured AxilJS AI provider, while vectorStore stores the vectors used during document retrieval.
Index Documents
Before querying the RAG pipeline, index your documents.
Use loadFromText() to create documents from text and pass them to rag.index().
Indexing prepares the document content for retrieval by the RAG pipeline.
For larger knowledge bases, the same indexing concept can be applied to a collection of documents.
Query the RAG Pipeline
Use rag.query() to ask a question against the indexed knowledge base.
The result contains both the generated answer and the sources used to produce it.
result.responsecontains the AI-generated answer.result.sourcescontains the documents retrieved and used as context.
Returning sources is useful when an application needs to show users where an answer came from.
Stream RAG Responses
AxilJS also supports streaming RAG queries with queryStream().
The streaming API exposes events as the RAG operation progresses.
The example handles two event types:
sources— Provides the retrieved source information.delta— Provides incremental generated response content.
This is useful for interactive RAG applications where the client needs to display retrieved sources and AI-generated content progressively.
Query vs Streaming
Use query() when your application needs the completed RAG result.
Use queryStream() when your application needs incremental events while the RAG operation is running.
| API | Output | Typical use |
|---|---|---|
rag.query() | Complete result | Standard RAG requests |
rag.queryStream() | Incremental events | Interactive RAG interfaces |
RAG Sources
RAG applications are often used when answers should be grounded in a defined document collection.
AxilJS exposes retrieved documents through result.sources for regular queries and source events during streaming.
This makes it possible to build interfaces that display both:
- The generated AI response
- The documents or sources used to generate it
Typical RAG Architecture
A production RAG application generally follows an indexing and query lifecycle:
The separation between indexing and querying allows documents to be prepared before users begin asking questions.
Common RAG Use Cases
The AxilJS RAG pipeline can be used for applications such as:
- Document question answering
- Knowledge-base assistants
- Technical documentation search
- Semantic document retrieval
- Internal knowledge systems
- AI-powered search
- Context-aware AI assistants
- Retrieval-based chat applications
RAG Workflow
A basic AxilJS RAG implementation follows these steps:
- Configure an AI provider.
- Create a
VectorStore. - Create a
RAGPipeline. - Load documents with
loadFromText(). - Index the documents with
rag.index(). - Query the indexed content with
rag.query(). - Use
result.responsefor the generated answer. - Use
result.sourcesto expose the retrieved sources. - Use
rag.queryStream()when streaming is required.