
Building Local AI Vector Search in .NET
Author - Abdul Rahman (Bhai)
AI
8 Articles
Table of Contents
What we gonna do?
A search box that only finds the exact words a reader typed makes useful content surprisingly hard to discover. Vector search, also called semantic search, solves that problem by comparing meaning rather than spelling. A query such as “how do I make my app search by meaning?” can therefore find content about embeddings even when it never uses that exact sentence.
This site replaces traditional keyword matching with a small, local-first AI vector search pipeline. It turns each article’s title, description, keywords, and channel into a 384-dimensional embedding at build time, then ranks the static vectors in the reader’s browser when they search. The result is a semantic search experience without a search API, server-side query processing, or a vector database.
The internet shows sustained interest in what is vector search, semantic search, vector database, and RAG (retrieval-augmented generation). Those terms describe a broad AI retrieval landscape; this article focuses on the deliberately smaller and more controllable static-site pattern used here.
Why we gonna do?
Keyword search fails whenever a reader and an author choose different words for the same idea. A title about “cancellation while accessing APIs” may be the right result for “stop HTTP requests when the user leaves,” but a literal token comparison has little evidence that the phrases belong together.
This mismatch becomes costly as a learning site grows because tags become inconsistent and every new synonym demands another manual keyword. Sending each keystroke to a hosted search service also adds latency, operational cost, and a new place where reader queries can leave the device. For a static catalogue, adding a remote database merely to compare a few hundred documents is often a rather expensive way to solve a small problem.
A local vector index gives the browser enough semantic context to rank the existing catalogue itself. The following implementation creates compatible normalized vectors for content and queries, then uses their dot product as a similarity score. In short: it makes search more forgiving while keeping the content index under the site’s deployment control.
Understand what “local” protects and what it does not
The generated index.json and index.vec files are shipped as static assets, so article vectors and ranking happen locally after download. The index also records the model identifier, immutable model revision, dimensions, entry count, and publish dates; browser code rejects incompatible metadata and invalid vector lengths before ranking.
However, “local index” is not automatically “fully offline model supply chain.” The current browser configuration permits the pinned embedding model to download remotely, while the build generator caches its pinned ONNX model and vocabulary locally after first use. That pin prevents an unreviewed latest-version change, but full offline control additionally requires serving reviewed model files from your own static assets and disabling remote model access.
How we gonna do?
Step 1: Define the search contract
Start with an interface so the Razor component does not know whether search uses keywords, vectors, or a remote provider. This repository keeps the contract deliberately small: warm the model and index, then search for article metadata.
namespace CommonComponents.Services;
// Keep the UI independent from the concrete keyword or vector-search implementation.
public interface IContentSearchService
{
// Warm the model and index before the first interactive search when possible.
Task WarmUpAsync(CancellationToken cancellationToken = default);
// Return article metadata so the UI never handles vector storage directly.
Task<IReadOnlyList<ContentMetaData>> SearchAsync(
string searchText,
CancellationToken cancellationToken = default);
}
Return ContentMetaData rather than raw vectors from this contract. The UI needs a title and URL, while the embedding dimensions and binary storage remain implementation details.
Step 2: Add a deterministic embedding index project
Create a console project that runs during the build and reference the project containing your content metadata. The generator in this repository uses ONNX Runtime for inference and Microsoft.ML.Tokenizers for BERT tokenization. Central package management keeps these versions controlled in Directory.Packages.props.
<Project Sdk="Microsoft.NET.Sdk">
<PropertyGroup>
<!-- Run this project as a build-time console tool. -->
<OutputType>Exe</OutputType>
<IsPackable>false</IsPackable>
</PropertyGroup>
<ItemGroup>
<!-- Read the shared article catalogue while generating vectors. -->
<ProjectReference Include="..\SharedModels\SharedModels.csproj" />
<!-- Run the pinned embedding model locally during the build. -->
<PackageReference Include="Microsoft.ML.OnnxRuntime" />
<PackageReference Include="Microsoft.ML.Tokenizers" />
</ItemGroup>
</Project>
Pin the model identity and revision in one place shared by the generator and browser module. This example uses Xenova/all-MiniLM-L6-v2, a MiniLM embedding model, revision 751bff37182d3f1213fa05d7196b954e230abad9, and 384 output dimensions.
Step 3: Convert metadata into normalized vectors
Build one searchable string from each article’s title, description, keywords, and channel. Normalize whitespace and symbols before tokenization so equivalent metadata produces stable input. The embedder truncates input to 256 tokens, mean-pools the token representations, and normalizes the result to unit length.
internal sealed class MiniLmEmbedder : IDisposable
{
public const int Dimensions = 384;
public const int MaxTokens = 256;
private readonly InferenceSession _session;
private readonly BertTokenizer _tokenizer;
public MiniLmEmbedder(string onnxModelPath, string vocabularyPath)
{
// Load the pinned ONNX model and its matching BERT vocabulary.
_session = new InferenceSession(onnxModelPath);
_tokenizer = BertTokenizer.Create(vocabularyPath, new BertOptions
{
LowerCaseBeforeTokenization = true,
ApplyBasicTokenization = true
});
}
public float[] Embed(string text)
{
// Convert normalized text into BERT token IDs.
var tokenIds = _tokenizer.EncodeToIds(text).ToArray();
if (tokenIds.Length > MaxTokens)
{
// Preserve the separator token after truncating long metadata.
tokenIds = tokenIds[..MaxTokens];
tokenIds[^1] = _tokenizer.SeparatorTokenId;
}
var length = tokenIds.Length;
// ONNX Runtime expects three BERT tensors with matching lengths.
var inputIds = new DenseTensor<long>([1, length]);
var attentionMask = new DenseTensor<long>([1, length]);
var tokenTypeIds = new DenseTensor<long>([1, length]);
for (var index = 0; index < length; index++)
{
inputIds[0, index] = tokenIds[index];
// Mark every token as active; token type IDs remain zero for one sequence.
attentionMask[0, index] = 1;
}
// Execute the pinned MiniLM ONNX model.
using var results = _session.Run([
NamedOnnxValue.CreateFromTensor("input_ids", inputIds),
NamedOnnxValue.CreateFromTensor("attention_mask", attentionMask),
NamedOnnxValue.CreateFromTensor("token_type_ids", tokenTypeIds)
]);
var hidden = results[0].AsTensor<float>();
var vector = new float[Dimensions];
// Mean-pool token states into one document embedding.
for (var token = 0; token < length; token++)
{
for (var dimension = 0; dimension < Dimensions; dimension++)
{
vector[dimension] += hidden[0, token, dimension];
}
}
for (var dimension = 0; dimension < Dimensions; dimension++)
{
vector[dimension] /= length;
}
var norm = MathF.Sqrt(vector.Sum(value => value * value));
// Unit-normalize so the dot product represents cosine similarity.
for (var dimension = 0; dimension < Dimensions; dimension++)
{
vector[dimension] /= norm;
}
return vector;
}
}
The shortened snippet highlights the algorithm; the complete implementation must populate all three tensors and perform the ONNX inference call. Do not replace mean pooling or normalization independently for documents and queries: both sides must use the same embedding procedure.
Create the document text from the same metadata that identifies the article in the catalogue. Including keywords and the channel gives the embedding model additional semantic signals without storing the article body in the vector index.
static string CreateSearchDocument(ContentMetaData content)
// Combine searchable metadata; article bodies are not stored in the vector index.
=> SearchTextNormalizer.Normalize(
$"{content.Title} {content.Description} " +
$"Keywords {string.Join(" ", content.Keywords)} " +
$"Channel {content.Channel}");
Step 4: Write metadata and binary vectors together
Store vector values as contiguous float32 values and store only the slug and publication date in JSON metadata. The array position is the join key: entry zero in index.json owns the first 384 floats in index.vec. Temporary files and final moves prevent a partially written index from becoming the deployed asset.
var contents = new TableOfContents().AllContents;
var entries = new List<SearchIndexEntry>(contents.Count);
// Keep vector order and metadata order identical: one slug owns 384 floats.
await using (var vectorStream = File.Create(temporaryVectorsPath))
{
foreach (var content in contents)
{
// Generate one normalized vector for the article metadata.
var vector = embedder.Embed(CreateSearchDocument(content));
vectorStream.Write(MemoryMarshal.AsBytes(vector.AsSpan()));
// Store only the lookup key and publication date in JSON metadata.
entries.Add(new SearchIndexEntry(
content.Slug,
content.ModifiedOn.ToString(
"yyyy-MM-dd",
CultureInfo.InvariantCulture)));
}
}
// Record the model contract so the browser can reject incompatible assets.
var index = new SearchIndexFile(
ModelId,
ModelRevision,
MiniLmEmbedder.Dimensions,
entries.Count,
entries);
// Serialize metadata separately from the compact binary vector file.
await File.WriteAllTextAsync(
temporaryMetadataPath,
JsonSerializer.Serialize(
index,
SearchIndexJsonContext.Default.SearchIndexFile));
// Replace final files only after generation succeeds.
File.Move(temporaryVectorsPath, vectorsPath, true);
File.Move(temporaryMetadataPath, metadataPath, true);
Step 5: Run the generator automatically from MSBuild
Add the generator source, project file, metadata sources, package versions, and browser module to the MSBuild input list. The target then reruns whenever one of those inputs changes and includes the generated files as static web assets.
<ItemGroup>
<!-- Rebuild when catalogue, generator, package, or browser code changes. -->
<SearchIndexInput Include="$(MSBuildThisFileDirectory)..\SharedModels\**\*.cs" />
<SearchIndexInput Include="$(MSBuildThisFileDirectory)..\VectorSearchIndexGenerator\**\*.cs"
Exclude="$(MSBuildThisFileDirectory)..\VectorSearchIndexGenerator\bin\**;$(MSBuildThisFileDirectory)..\VectorSearchIndexGenerator\obj\**" />
<SearchIndexInput Include="$(MSBuildThisFileDirectory)..\VectorSearchIndexGenerator\VectorSearchIndexGenerator.csproj" />
<SearchIndexInput Include="$(MSBuildThisFileDirectory)..\Directory.Packages.props" />
<SearchIndexInput Include="wwwroot\js\vector-search.js" />
</ItemGroup>
<Target Name="GenerateContentVectorIndex"
BeforeTargets="ResolveProjectStaticWebAssets"
Inputs="@(SearchIndexInput)"
Outputs="wwwroot\search\index.json;wwwroot\search\index.vec">
<!-- Generate the index before static web assets are resolved. -->
<Exec Command="dotnet run --project "$(MSBuildThisFileDirectory)..\VectorSearchIndexGenerator\VectorSearchIndexGenerator.csproj" -- "$(MSBuildThisFileDirectory)wwwroot\search"" />
</Target>
Run dotnet build Web/Web.csproj after adding or editing content. Inspect CommonComponents/wwwroot/search/index.json to confirm the article slug and model revision, and keep the binary index beside it.
Step 6: Load and validate the index in the browser
Import the repository’s bundled Transformers.js module and cache both the model promise and index promise. Validate the model identity, revision, dimensions, entry count, dates, and vector length before calculating a score. This fails fast instead of silently ranking vectors generated by a different model.
import { env, pipeline } from '../lib/transformers/transformers.min.js';
// These values must match the build-time generator exactly.
const dimensions = 384;
const model = 'Xenova/all-MiniLM-L6-v2';
const modelRevision = '751bff37182d3f1213fa05d7196b954e230abad9';
// The current repository allows the pinned model to be fetched remotely.
env.allowLocalModels = false;
env.allowRemoteModels = true;
async function initializeIndex() {
// Load metadata and binary vectors together from static web assets.
const [metadataResponse, vectorResponse] = await Promise.all([
fetch(new URL('search/index.json', document.baseURI)),
fetch(new URL('search/index.vec', document.baseURI)),
]);
if (!metadataResponse.ok || !vectorResponse.ok) {
throw new Error('The vector search index could not be loaded.');
}
const index = await metadataResponse.json();
const vectors = new Float32Array(
await vectorResponse.arrayBuffer());
// Reject assets generated with another model or incompatible dimensions.
if (index.Model !== model
|| index.ModelRevision !== modelRevision
|| index.Dimensions !== dimensions
|| !Array.isArray(index.Entries)
|| index.Count !== index.Entries.length
|| vectors.length !== index.Count * dimensions) {
throw new Error('The vector index is incompatible.');
}
return { entries: index.Entries, vectors };
}
Step 7: Embed the query and rank the top results
Embed the normalized query with the same model and options used for the document vectors. Filter future publication dates, calculate the dot product against each 384-value slice, discard scores below 0.25, sort descending, and return the requested maximum number of slugs.
// Embed the query with the same pooling and normalization as documents.
const output = await extractor(
normalizedText,
{ pooling: 'mean', normalize: true });
const queryVector = output.data;
// Rank published entries, apply the threshold, and cap the result count.
return entries
.map((entry, entryIndex) => ({ entry, entryIndex }))
.filter(candidate => candidate.entry.PublishOn <= localDate())
.map(candidate => ({
slug: candidate.entry.Slug,
score: dotProduct(
queryVector,
vectors,
candidate.entryIndex),
}))
.filter(candidate => candidate.score >= 0.25)
.sort((left, right) => right.score - left.score)
.slice(0, maximumResults)
.map(result => result.slug);
Query: "find articles about semantic retrieval"
Candidate article Similarity Outcome
--------------------------------------- ---------- ----------------
Local AI vector search in .NET 0.82 show first
Building an AI chat application 0.43 show later
Introduction to logging 0.11 below threshold
Step 8: Bridge JavaScript results back to .NET
Keep JavaScript responsible for numerical work and return only slugs through JavaScript interop. The .NET service resolves those slugs against TableOfContents.AllContents, applies a second future-date check, and returns complete ContentMetaData objects to the component.
public async Task<IReadOnlyList<ContentMetaData>> SearchAsync(
string searchText,
CancellationToken cancellationToken = default)
{
// Avoid loading the model for an empty query.
if (string.IsNullOrWhiteSpace(searchText))
{
return [];
}
// Call the cached browser module and request at most ten slugs.
var module = await GetModuleAsync();
var slugs = await module.InvokeAsync<string[]>(
"search",
cancellationToken,
searchText,
10);
// Resolve slugs to trusted catalogue metadata and hide scheduled content.
return [
.. slugs
.Select(slug => _contentsBySlug.GetValueOrDefault(slug))
.Where(content => content is not null
&& content.ModifiedOn.Date <= DateTime.Today)
.Cast<ContentMetaData>()
];
}
Register the implementation in both application entry points so Blazor WebAssembly and .NET MAUI use the same service. The existing applications do this with a scoped lifetime because the JavaScript module and content lookup belong to one app session.
// Use the same implementation in both WebAssembly and MAUI entry points.
services.AddScoped<IContentSearchService, VectorContentSearchService>();
Connect the search box without racing requests
In Search.razor.cs, debounce input for 400 milliseconds, cancel the previous token, and increment a search version. Only the latest version may update the suggestions list; this prevents a slow query from replacing results for newer text. Warm the model after the first render, or earlier when the input receives focus.
// Debounce keystrokes so every character does not start an embedding operation.
await Task.Delay(400, cancellationToken);
var results = await ContentSearchService.SearchAsync(
searchText,
cancellationToken);
// Ignore an older query that completed after newer input arrived.
if (searchVersion != _searchVersion
|| !searchText.Equals(
SearchText,
StringComparison.Ordinal))
{
return;
}
_filteredContents = [.. results];
ShowSuggestions = _filteredContents.Count > 0;
BUILD TIME
TableOfContents --> VectorSearchIndexGenerator
| |
| +-- pinned ONNX model + vocabulary
v v
article metadata --> index.json + index.vec --> static deployment
SEARCH TIME
Search.razor.cs --> IContentSearchService --> vector-search.js
^ |
| +-- query embedding
| +-- local index loading
| +-- score, filter, sort
+----------- article slugs <------------------+
Harden the local supply chain before calling it fully offline
The current repository pins the model revision and package lockfile, but its browser module sets allowRemoteModels to true. For a fully local runtime, vendor the reviewed model files under the application’s static assets, set allowLocalModels to true, set allowRemoteModels to false, verify hashes in CI, and test with the model host blocked. Keep the build-time model cache and browser model revision identical so document and query vectors remain compatible.
env.allowLocalModels = true;
env.allowRemoteModels = false;
// Point Transformers.js at model files reviewed and deployed by your application.
env.localModelPath = new URL(
'./models/',
document.baseURI).toString();
Exact scoring scans every vector, which is appropriate for this site’s modest catalogue. If the catalogue becomes large, replace the linear scan with an approximate-nearest-neighbour index or a dedicated vector store, but regenerate every document vector whenever the embedding model changes.
Summary
Local AI vector search gives a static .NET site a useful semantic retrieval layer without sending every query to a search API or server-side query processor. It works because documents and queries use the same embedding model, normalization makes scores comparable, and the browser can rank the published static index directly.
- Embeddings represent the meaning of content as numeric vectors rather than literal keywords.
- Normalized dot products rank semantic similarity and let a threshold discard weak matches.
- Build-time indexing keeps search data static, deployable, and independent of a vector database.
- Revision pinning and validation protect compatibility, but locally hosting model assets is required for fully offline model delivery.
- Debouncing, cancellation, and publish-date checks keep the reader experience responsive and correct.
What to read next: explore AI chat applications in .NET with Microsoft.Extensions.AI or learn how GitHub Copilot can help you navigate a new codebase.