AI Platform

AxilJS RAG Pipeline for Retrieval-Augmented Generation

Build retrieval-augmented generation (RAG) applications with AxilJS using document indexing, vector search, AI-generated answers, source retrieval, and streaming.

4 min readDocumentationEdit this page

AxilJS RAG Pipeline

AxilJS provides a Retrieval-Augmented Generation (RAG) pipeline for building AI applications that retrieve relevant information from indexed documents and use that information to generate contextual answers.

The RAGPipeline combines an AI provider with a VectorStore, allowing applications to index text, perform retrieval-based queries, return source documents, and stream RAG responses.

What Is RAG?

Retrieval-Augmented Generation combines two stages:

  1. Retrieval — Find relevant information from indexed documents using vector search.
  2. Generation — Pass the retrieved information to an AI provider to generate an answer.

A typical AxilJS RAG workflow looks like this:

text
Documents
    │
    ▼
loadFromText()
    │
    ▼
RAGPipeline.index()
    │
    ▼
VectorStore
    │
    │  User Query
    ▼
RAGPipeline.query()
    │
    ├──► Relevant Sources
    │
    ▼
AI Provider
    │
    ▼
Generated Response

This approach allows an AI application to answer questions using information retrieved from a specific document collection.

Setup

Import RAGPipeline, VectorStore, and loadFromText from @axiljs/ai.

typescript
import {
  RAGPipeline,
  VectorStore,
  loadFromText
} from '@axiljs/ai'
 
const rag = new RAGPipeline({
  provider: ai,
  vectorStore: new VectorStore()
})

The provider is the configured AxilJS AI provider, while vectorStore stores the vectors used during document retrieval.

Index Documents

Before querying the RAG pipeline, index your documents.

Use loadFromText() to create documents from text and pass them to rag.index().

typescript
await rag.index(
  loadFromText([
    'AxilJS supports PostgreSQL...'
  ])
)

Indexing prepares the document content for retrieval by the RAG pipeline.

For larger knowledge bases, the same indexing concept can be applied to a collection of documents.

Query the RAG Pipeline

Use rag.query() to ask a question against the indexed knowledge base.

typescript
const result = await rag.query(
  'What databases are supported?'
)
 
result.response  // generated answer
result.sources   // documents used

The result contains both the generated answer and the sources used to produce it.

  • result.response contains the AI-generated answer.
  • result.sources contains the documents retrieved and used as context.

Returning sources is useful when an application needs to show users where an answer came from.

Stream RAG Responses

AxilJS also supports streaming RAG queries with queryStream().

typescript
for await (
  const event of rag.queryStream('Tell me about auth')
) {
  if (event.type === 'sources') {
    console.log(event.data)
  }
 
  if (event.type === 'delta') {
    process.stdout.write(event.data)
  }
}

The streaming API exposes events as the RAG operation progresses.

The example handles two event types:

  • sources — Provides the retrieved source information.
  • delta — Provides incremental generated response content.

This is useful for interactive RAG applications where the client needs to display retrieved sources and AI-generated content progressively.

Query vs Streaming

Use query() when your application needs the completed RAG result.

typescript
const result = await rag.query(
  'What databases are supported?'
)
 
console.log(result.response)

Use queryStream() when your application needs incremental events while the RAG operation is running.

typescript
for await (
  const event of rag.queryStream('Tell me about auth')
) {
  // Process each event
}
APIOutputTypical use
rag.query()Complete resultStandard RAG requests
rag.queryStream()Incremental eventsInteractive RAG interfaces

RAG Sources

RAG applications are often used when answers should be grounded in a defined document collection.

AxilJS exposes retrieved documents through result.sources for regular queries and source events during streaming.

This makes it possible to build interfaces that display both:

  • The generated AI response
  • The documents or sources used to generate it

Typical RAG Architecture

A production RAG application generally follows an indexing and query lifecycle:

text
             INDEXING
                 │
                 ▼
          Source Documents
                 │
                 ▼
          loadFromText()
                 │
                 ▼
          RAGPipeline.index()
                 │
                 ▼
            VectorStore
                 │
                 │
                 ▼
              QUERY
                 │
          User Question
                 │
                 ▼
          RAGPipeline.query()
                 │
        ┌────────┴────────┐
        ▼                 ▼
  Relevant Sources    AI Provider
        │                 │
        └────────┬────────┘
                 ▼
          Generated Answer

The separation between indexing and querying allows documents to be prepared before users begin asking questions.

Common RAG Use Cases

The AxilJS RAG pipeline can be used for applications such as:

  • Document question answering
  • Knowledge-base assistants
  • Technical documentation search
  • Semantic document retrieval
  • Internal knowledge systems
  • AI-powered search
  • Context-aware AI assistants
  • Retrieval-based chat applications

RAG Workflow

A basic AxilJS RAG implementation follows these steps:

  1. Configure an AI provider.
  2. Create a VectorStore.
  3. Create a RAGPipeline.
  4. Load documents with loadFromText().
  5. Index the documents with rag.index().
  6. Query the indexed content with rag.query().
  7. Use result.response for the generated answer.
  8. Use result.sources to expose the retrieved sources.
  9. Use rag.queryStream() when streaming is required.

Help improve the documentation

AxilJS is open source and documentation improvements are welcome.

AxilJS DocumentationMIT License · Built by SyntaxilitY