A language model can only answer from what it saw in training, so an assistant that has to cite a firm's documents or this week's data needs those sources in the prompt. Retrieval-augmented generation (RAG) is the standard architecture for LLM-powered applications that do this: it fetches relevant text from an external knowledge base and hands it to the model at query time, which improves accuracy and lets answers include current or proprietary data.
How RAG works
A RAG system runs in two steps:
- Retrieval fetches relevant information from a knowledge base. Documents are chunked and converted into vector embeddings with a model like OpenAI's
text-embedding-3-smallor an open-source option likesentence-transformers, and the embeddings are stored in a vector database. - Generation embeds the incoming query into the same vector space, runs a similarity search for the closest chunks, and passes them as context to a language model like GPT-4, which writes the answer.
Because the model and the data are decoupled, you can change the knowledge base without retraining anything, which suits domain-specific applications whose sources change often.
Similarity search
Semantic search is the core of retrieval. Instead of matching exact words, embeddings place text in a high-dimensional space where proximity reflects meaning, so "attorney" and "lawyer" land close together even when a document uses only one of them.
Similarity is measured with a distance metric such as cosine similarity, and vector databases use approximate nearest neighbor (ANN) algorithms to find the closest chunks without comparing every vector. The result is retrieval that matches a query to documents by concept rather than by syntax.
Where RAG fits
RAG is the right architecture for:
- Domain-specific assistants that answer from a custom knowledge base.
- Chatbots that have to stay accurate as their sources change.
- Document Q&A across PDFs, websites, or internal wikis.
Node.js in the RAG stack
Node.js is a practical choice for RAG when a team already works in JavaScript and TypeScript, because the embedding APIs and integrations are mature enough to cover the whole pipeline:
- Generate embeddings with OpenAI or the Hugging Face Inference API.
- Store and query vectors in
pgvector, Pinecone, Weaviate, or Qdrant. - Run end-to-end workflows in Next.js server functions, Vercel edge functions, or Node.js servers.
A growing set of open-source SDKs built for LLM apps covers the rest, with patterns for streaming responses, function calling, and vector search.
Open-source TypeScript AI SDKs
Vercel AI SDK
Vercel's TypeScript-first toolkit for AI apps, with streaming and tool use across React, Next.js, and Node.js:
- A unified API for OpenAI, Anthropic, and other providers.
- Streaming UI and edge-runtime compatibility.
- Tool calling with typed context.
- https://github.com/vercel/ai
LangChain.js
The JavaScript and TypeScript version of LangChain, for LLM chains, agents, and RAG pipelines:
- Chains, tools, agents, retrievers, and memory.
- Integrations with vector stores like Pinecone, Supabase, and Weaviate.
- Structured output parsing, evals, and streaming.
- https://github.com/langchain-ai/langchainjs
LlamaIndex.TS
The TypeScript version of LlamaIndex, for ingesting, indexing, and querying data in RAG systems:
- Indexing for documents and metadata.
- LLM-driven query planning.
- Retriever composition and routing.
- https://github.com/jerryjliu/llamaindex-ts
Mastra
A TypeScript agent framework with persistent memory, threads, and embedded knowledge graphs:
- Agents with tool use and long-term memory.
- JSON knowledge graphs and fact-based reasoning.
- Context-aware multi-step workflows.
- https://github.com/mastra-ai/mastra
Agentic.so
A standard library of TypeScript AI tools that exposes functions to LLMs cleanly:
- Typed agent workflows.
- AI-accessible function definitions.
- An SDK for tool abstraction.
- https://github.com/agentic-dev/agentic
Vector storage options
Where the vectors live depends on the stack you already run:
pgvectoris a PostgreSQL extension for vector similarity search, and it fits full-stack apps that already use Postgres.- Dedicated vector databases like Pinecone, Weaviate, and Qdrant add scalable nearest-neighbor search and metadata filtering.
- OpenAI's managed vector store, released in 2024, is integrated with their embeddings and models and suits small-to-medium workloads where you want little operational overhead.
RAG in projects I've shipped
On LegalAgent, RAG supplies case context, document summaries, and procedural guidance to a legal assistant, and on Bitlauncher, an AI and crypto launchpad, the chatbot uses RAG for longer answers about projects. Python still leads the AI tooling ecosystem, but the TypeScript stack has been enough for both, and it lets web engineers ship AI features without switching languages.