Ingestion and embeddings

Pipelines

Ingestion and embeddings

When a trainer uploads a resource the API stores the file, writes a row and queues a job. A worker in the Python ai-service extracts the text, splits it into chunks of about 500 tokens, asks Ollama for a 768-dimension vector per chunk, and stores the results in resource_chunks.

Sequence

Ingestion sequence

Ollamaai-service workerPostgresVolume uploads/Express APIloop[poll every few seconds]loop[each chunk]TrainerPOST /api/resources (multipart)1saveUpload() writes the file2INSERT learning_resources, enqueueEmbedJob()3201 Created4claim_job() FOR UPDATE SKIPLOCKED5read file6extract_text(), chunk_text()7embed() POST /api/embeddings8768 floats9store_chunks() as super_admin10mark_done()11
The upload request returns as soon as the job is queued. Embedding happens afterwards.

Job lifecycle

Flow of one job

exception

Trainer uploads resource

Job queue
Postgres table

Worker claims a job
app role: super_admin

extract_text and chunk_text

embed each chunk
Ollama nomic-embed-text

resource_chunks
VECTOR 768

mark_done

mark_failed, retry with backoff

A worker claims a job, embeds each chunk, stores the vectors and marks the job done. A failure schedules a retry.

Functions

The API side is short. Almost everything lives in the worker.

FunctionRuns inDoes
saveUpload(file, courseId)APIValidates type and size, writes to uploads/<course>/<uuid>.<ext>, returns the relative path.
createResource(req)APIRoute handler for POST /api/resources. In one transaction: INSERT learning_resources, then enqueueEmbedJob.
enqueueEmbedJob(db, resourceId)APIINSERT into the queue table with status queued.
claim_job(conn)ai-serviceAtomically takes one queued job using FOR UPDATE SKIP LOCKED.
extract_text(path, mime)ai-serviceText from pptx, pdf or transcript files.
chunk_text(text, target_tokens=500, overlap=50)ai-serviceSplits on paragraph, then sentence boundaries.
embed(text)ai-servicePOST to Ollama, checks the vector has 768 numbers.
store_chunks(conn, resource_id, chunks, vectors)ai-serviceDeletes old chunks for the resource, inserts the new ones, in one transaction.
mark_done(conn, job_id)ai-servicestatus done. On an exception, mark_failed sets status failed and schedules a retry.

Worker sketch

services/ai-service/worker.py
import os, requests

OLLAMA = os.environ["OLLAMA_HOST"]                    # never hardcode localhost
MODEL  = os.environ.get("EMBED_MODEL", "nomic-embed-text")

def embed(text: str) -> list[float]:
    r = requests.post(f"{OLLAMA}/api/embeddings",
                      json={"model": MODEL, "prompt": text}, timeout=60)
    r.raise_for_status()
    v = r.json()["embedding"]
    assert len(v) == 768, f"expected 768 dims, got {len(v)}"
    return v

def process(conn, job):
    path, mime = load_resource_row(conn, job["resource_id"])
    text   = extract_text(path, mime)
    chunks = chunk_text(text, target_tokens=500, overlap=50)
    vecs   = [embed(c) for c in chunks]
    with conn.transaction():
        set_role(conn, "super_admin")                  # set_config('app.current_role', ...)
        store_chunks(conn, job["resource_id"], chunks, vecs)   # delete old, insert new
        mark_done(conn, job["id"])
claim query
UPDATE embed_jobs
   SET status = 'running', started_at = now()
 WHERE id = (SELECT id FROM embed_jobs
              WHERE status = 'queued'
              ORDER BY created_at
              FOR UPDATE SKIP LOCKED
              LIMIT 1)
RETURNING id, resource_id;

Rules and gotchas

Idempotent store

store_chunks deletes the resource's old chunks and inserts new ones in one transaction. A retry after a crash cannot leave duplicates.

Stuck jobs

Workers can die mid-job. A periodic sweep should return jobs that stayed running past a timeout to queued.

Model name travels with the vector

Write embedding_model on every chunk. Query time should only compare vectors from the same model.

Prompt prefixes

Nomic embedding models were trained with task prefixes such as search_document: and search_query:. Pick one convention, apply it everywhere, and record it in embedding_model, for example nomic-embed-text/search_document.

Batch endpoint

Newer Ollama versions accept a list of inputs at /api/embed and return embeddings. Use it if the installed version supports it, otherwise keep one call per chunk.

Capacity Connect · Team Syntax Squad · SIH 2026 · PS 26075Code samples are implementation sketches.