Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 7 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
@@ -1,6 +1,11 @@
# Environment files
.env

# Runtime config (secrets) — copy config.yaml.example → config.yaml and fill in
config.yaml
config.yaml.bak*
*.bak*

# Python
__pycache__/
*.py[cod]
Expand Down Expand Up @@ -34,10 +39,11 @@ env/
*.swp
*.swo

# Model cache
# Model cache / embedding cache
.cache/
models/
sentence_transformers/
embed_cache.json

# Testing
.pytest_cache/
Expand Down
214 changes: 214 additions & 0 deletions SETUP.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,214 @@
# Vector Indexer MCP — Deployment Setup Guide

This document describes the **actual running deployment** — the `*_local.py` modules,
migrations, and systemd units that ship in this repo. The original `src/server.py` /
`daemon/*.py` modules (Supabase backend) remain in tree but are not used by the
production deployment.

---

## Prerequisites

- Python 3.10+
- PostgreSQL 14+ with **pgvector** extension
- systemd (user session) for service management

---

## 1. PostgreSQL / pgvector

### Install pgvector (if not already installed)

```bash
# Debian/Ubuntu
sudo apt install postgresql-14-pgvector # or postgresql-15-pgvector

# Build from source (any distro)
git clone https://github.com/pgvector/pgvector.git
cd pgvector && make && sudo make install
```

### Create role and database

```sql
-- as postgres superuser
CREATE ROLE vectoruser WITH LOGIN PASSWORD 'your-strong-password';
CREATE DATABASE vectorindex OWNER vectoruser;

-- connect to vectorindex and enable the extension
\c vectorindex
CREATE EXTENSION IF NOT EXISTS vector;
GRANT ALL ON SCHEMA public TO vectoruser;
```

### Apply the migration

```bash
psql -U vectoruser -d vectorindex -f migrations/020_vector_search_indexer.sql
```

This creates:
- `file_metadata`, `file_chunks`, `file_embeddings`, `index_queue`, `index_stats` tables
- GIN index for full-text search
- HNSW index for vector cosine similarity
- `search_hybrid()` and `get_index_health()` stored functions

---

## 2. Python environment

```bash
cd /path/to/vector-indexer-mcp

python3 -m venv venv
source venv/bin/activate

pip install -r requirements.txt
```

Key runtime dependencies (all in `requirements.txt`):
- `asyncpg` — PostgreSQL driver used by all `*_local.py` modules
- `sentence-transformers` + `torch` — embedding generation
- `watchdog` — file system monitoring
- `mcp` — MCP server protocol
- `uvicorn`, `starlette`, `sse-starlette` — HTTP/SSE transport (`sse_server_local.py`)

> **Note on pyproject.toml**: The `[project.scripts]` entries (`vector-indexer-worker`,
> `vector-indexer-mcp`) point at the old Supabase modules (`daemon.worker:main`,
> `src.server:main`). Do **not** use them for the local deployment — use the systemd
> units or the module invocations below instead.

---

## 3. Configuration

```bash
cp config.yaml.example config.yaml
# Edit config.yaml and set:
# database.password → your PostgreSQL password (same as step 1)
# watcher.paths → list of absolute directory paths to monitor
```

**Important**: `config.yaml` is gitignored and must never be committed — it contains
your database password. The modules load it from a **hardcoded path**:

```
/home/david/vector-indexer-mcp/config.yaml # in src/server_local.py
/home/david/vector-indexer-mcp/config.yaml # in daemon/worker_local.py
/home/david/vector-indexer-mcp/config.yaml # in daemon/watcher_local.py
```

This path is hardcoded in the source (known limitation — see Recommended Follow-ups).
For a different checkout location, update `CONFIG_PATH` in each `*_local.py` file or
create a symlink.

### watcher.paths gotcha

The `watcher.paths` list accepts **any** absolute path, including paths outside the
project root. The daemon watches each listed directory recursively. Paths that do not
exist at startup are skipped with a warning (they can be added later by restarting the
service).

---

## 4. Systemd user services

The `systemd/` directory contains the three unit files. Copy them to your user systemd
directory and adjust the `WorkingDirectory` and `ExecStart` paths to match your
checkout location before enabling.

```bash
# Adjust paths inside each .service file first, then:
cp systemd/vector-indexer-worker.service ~/.config/systemd/user/
cp systemd/vector-indexer-watcher.service ~/.config/systemd/user/
cp systemd/vector-indexer-mcp-http.service ~/.config/systemd/user/

systemctl --user daemon-reload

# Start worker first (embedder / queue processor)
systemctl --user enable --now vector-indexer-worker

# Start file watcher (enqueues changed files)
systemctl --user enable --now vector-indexer-watcher

# Start HTTP/SSE server (optional — for non-stdio MCP clients)
systemctl --user enable --now vector-indexer-mcp-http
```

Check status / logs:

```bash
systemctl --user status vector-indexer-worker
journalctl --user -u vector-indexer-worker -f
```

### Service startup order

`vector-indexer-worker` → `vector-indexer-mcp-http` (declared in `After=`).
`vector-indexer-watcher` is independent of the other two.

---

## 5. Claude MCP stdio entry

For direct stdio usage (the typical Claude Code / claude-mpm setup), add this to your
`mcpServers` configuration:

```json
{
"vector-indexer-mcp": {
"command": "/path/to/vector-indexer-mcp/venv/bin/python",
"args": ["-m", "src.server_local"],
"cwd": "/path/to/vector-indexer-mcp"
}
}
```

The server reads `config.yaml` at import time (hardcoded path — see note above), so
the `cwd` must match the checkout location, **or** the `CONFIG_PATH` constant in
`src/server_local.py` must be updated.

---

## 6. Initial bulk index

After the services are running, queue all files in a watched directory for indexing:

```bash
# Via MCP tool (from a Claude session):
# search_hybrid → reindex_path(path="/your/project", recursive=true, force=true)

# Or directly via psql:
INSERT INTO index_queue(file_path, event_type, status)
SELECT file_path, 'modify', 'pending'
FROM file_metadata;
```

Or use `scripts/bulk_index.py` if it exists.

---

## 7. Verify the index

```bash
source venv/bin/activate
python scripts/verify_index.py
```

Or via MCP `index_status` tool — returns `total_files`, `total_chunks`,
`total_embeddings`, and `queue.pending`.

---

## Recommended Follow-ups (not done in this reconciliation)

1. **De-hardcode `CONFIG_PATH`** — read from `VECTOR_INDEXER_CONFIG` env var with a
fallback, so the service works from any checkout location without editing source.
2. **Fix `pyproject.toml` scripts** — `vector-indexer-worker` and `vector-indexer-mcp`
point at the old Supabase modules; update them to point at `daemon.worker_local:main`
and `src.server_local:main` (or remove them and rely on the systemd units).
3. **Prune dead dependencies** — `supabase`, `postgrest-py`, and `tiktoken` are listed
in `requirements.txt` / `pyproject.toml` but not used by the `*_local.py` runtime.
Removing them reduces install size and eliminates a large dependency surface.
4. **Add `asyncpg` to `pyproject.toml` dependencies** — currently only in
`requirements.txt`; the two should stay in sync.
35 changes: 35 additions & 0 deletions config.yaml.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
database:
host: localhost
port: 5432
name: vectorindex
user: vectoruser
# Replace with your actual database password — never commit the real value
password: CHANGE_ME

watcher:
# List the absolute paths you want the watcher to monitor.
# These are machine-specific; adjust for your environment.
# Note: paths outside the project root (e.g. /home/user/other-repo) are
# supported — the watcher follows them correctly.
paths:
- /path/to/your/project-one
- /path/to/your/project-two
debounce_ms: 500
extensions: [.py, .md, .txt, .yaml, .yml, .json, .toml, .sh, .sql, .ts, .js, .tsx, .jsx]
exclude_patterns: [__pycache__, .git, node_modules, .venv, venv, .cache, dist, build, .next, __snapshots__]

embedding:
model: paraphrase-multilingual-MiniLM-L12-v2
batch_size: 32
device: cpu # set to "cuda" if a GPU is available

chunking:
chunk_size_tokens: 500
overlap_tokens: 50

queue:
poll_interval_ms: 500
batch_size: 50

server:
port: 5577
82 changes: 82 additions & 0 deletions daemon/chunker_local.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,82 @@
"""Token-aware text chunker for local asyncpg-based vector indexer."""

from dataclasses import dataclass
from typing import List


@dataclass
class TextChunk:
text: str
start_line: int
end_line: int
chunk_index: int
token_count: int


class TextChunker:
def __init__(self, chunk_size_tokens: int = 500, overlap_tokens: int = 50):
self.chunk_size = chunk_size_tokens
self.overlap = overlap_tokens
try:
import tiktoken

self.enc = tiktoken.get_encoding("cl100k_base")
self._use_tiktoken = True
except Exception:
self._use_tiktoken = False

def _count_tokens(self, text: str) -> int:
if self._use_tiktoken:
return len(self.enc.encode(text))
return len(text.split())

def chunk_text(self, text: str, file_path: str = "") -> List[TextChunk]:
lines = text.splitlines(keepends=True)
chunks = []
current_lines = []
current_tokens = 0
chunk_idx = 0
start_line = 1

for line_num, line in enumerate(lines, 1):
line_tokens = self._count_tokens(line)
if current_tokens + line_tokens > self.chunk_size and current_lines:
chunk_text = "".join(current_lines)
chunks.append(
TextChunk(
text=chunk_text,
start_line=start_line,
end_line=line_num - 1,
chunk_index=chunk_idx,
token_count=current_tokens,
)
)
chunk_idx += 1
# Overlap: keep last N tokens worth of lines
overlap_lines = []
overlap_tokens = 0
for prev_line in reversed(current_lines):
t = self._count_tokens(prev_line)
if overlap_tokens + t > self.overlap:
break
overlap_lines.insert(0, prev_line)
overlap_tokens += t
current_lines = overlap_lines + [line]
current_tokens = overlap_tokens + line_tokens
start_line = line_num - len(overlap_lines)
else:
current_lines.append(line)
current_tokens += line_tokens

if current_lines:
chunks.append(
TextChunk(
text="".join(current_lines),
start_line=start_line,
end_line=len(lines),
chunk_index=chunk_idx,
token_count=current_tokens,
)
)

return chunks
29 changes: 29 additions & 0 deletions daemon/embedder_local.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
"""Embedding generator for local asyncpg-based vector indexer."""

from typing import List


class EmbeddingGenerator:
def __init__(
self,
model_name: str = "paraphrase-multilingual-MiniLM-L12-v2",
device: str = "cpu",
):
self.model_name = model_name
self.device = device
self._model = None

def _get_model(self):
if self._model is None:
from sentence_transformers import SentenceTransformer

self._model = SentenceTransformer(self.model_name, device=self.device)
return self._model

def embed_batch(self, texts: List[str]) -> List[List[float]]:
model = self._get_model()
embeddings = model.encode(texts, batch_size=32, show_progress_bar=False)
return [e.tolist() for e in embeddings]

def embed_single(self, text: str) -> List[float]:
return self.embed_batch([text])[0]
Loading