Back to blog
How to Build People Search With a Knowledge Graph

This blog is written by AI for SEO

How to Build People Search With a Knowledge Graph

HelixDB9 min read

A recruiter types "founding engineer who scaled distributed storage at a seed-stage startup and worked with Kubernetes internals." Pure vector search over raw resume text fails this query immediately. The embedding model matches stylistic resumes full of buzzwords, surfacing lookalikes who wrote documentation at large enterprises while missing the candidates who actually built storage engines at small startups.

Semantic similarity is blind to relational constraints. A vector space can't enforce that Person A worked at Company B during Year C, or that Person A reported to Person D on Project E. Accurate candidate retrieval needs both graph topology and semantic vector similarity.

This guide shows you how to build a people search knowledge graph from scratch. Before you write code, you need profile data, an embedding strategy, and an entity resolution pipeline. Then you model the data as a graph and use HelixDB to prefilter vector search directly on graph edges.

Step 1: Decide what profile data to keep (and what to drop)

Most teams building talent search dump raw LinkedIn exports or scraper output straight into an embedding pipeline. That creates noise. Unstructured profile text is full of fluff like "passionate problem solver" and "visionary leader," and that fluff pulls vector distance toward empty jargon.

Split profile data into two explicit categories: structured relational facts and contextual narrative blocks. Relational facts belong in graph properties and edges. Narratives belong in text chunks destined for embeddings and BM25 full-text indexing.

Keep these structured fields for your graph schema: full names, previous company identifiers, job titles, start and end dates, school names, degrees, and explicit project tags. Drop generic summary bios that contain no technical facts or metrics. Strip marketing filler from role descriptions before you do anything else.

Keep the narrative blocks that describe actual execution. "Built an internal raft consensus layer handling 40k writes per second" stays. "Collaborated cross-functionally across stakeholders to drive alignment" goes. Filtering out low-signal text cuts token spend, keeps your vectors focused on real competence, and removes semantic noise before indexing.

Step 2: Chunk bios and work history so each chunk answers one question

Don't embed an entire resume as a single 1500-token blob. When an embedding model compresses ten years of varied employment history into one 1536-dimension vector, specific achievements get blurred. A recruiter searching for someone who wrote custom eBPF probes will miss an engineer whose eBPF work got squashed together with six years of frontend React.

Chunk work history by discrete role and specific accomplishment. Each chunk must answer one question: what did this person achieve inside this organization or project? If an engineer spent three years at Datadog as a senior engineer and two years at Stripe as a staff engineer, create separate chunks for each tenure.

Each chunk should carry its subject context. Instead of embedding an isolated fragment like "Optimized p99 query latency from 450ms to 12ms," prepend the structure: "Role: Staff Backend Engineer at Stripe (2022 to 2024). Accomplishment: Optimized p99 query latency from 450ms to 12ms using custom RocksDB compaction filters."

Prepending anchors the vector to both the role and the domain without needing an oversized context window. When each chunk maps to a node or edge in your graph, your retriever can match exact technical accomplishments and still know where and when the work happened.

Step 3: Resolve people and companies across sources before you embed anything

Entity resolution makes or breaks a people search knowledge graph. If your raw data lists "Google," "Google LLC," and "Alphabet Inc." as separate entities, your graph loses connectivity. A query for colleagues who overlapped at Google returns empty traversals because the engineers hang off different company nodes.

Canonicalize company and school entities before computing embeddings or inserting graph edges. For companies, normalize names against official registry identifiers, domain names, or curated entity databases. Follow the entity resolution patterns from LinkedIn engineering, where every organization maps to a canonical node ID with known alternate aliases.

Person deduplication needs stricter heuristics. Combine normalized email hashes, personal portfolio URLs, GitHub handles, and overlapping employment dates. If two profiles share a full name but have conflicting work histories, keep them separate. Merging two distinct people into one profile corrupts the graph permanently. Missing an occasional duplicate is easy to fix later.

Entity linking research shows that resolving noisy entity names before indexing prevents graph fragmentation. Once you have canonical IDs for companies, universities, and open-source repositories, you can reliably connect candidate profiles across sources.

Step 4: Model Person, Company, Skill, and Project as a graph

Model your data as a property graph with four core node labels: Person, Company, Skill, and Project. Don't turn every resume keyword into a node. If generic terms like "communication" or "agile" become nodes, your graph grows dense supernodes that slow multi-hop traversals and add nothing to search.

Define explicit edge types. Use WORKED_AT between Person and Company, with properties like title, start_year, end_year, and is_current. Use HAS_SKILL between Person and Skill, with an integer years_experience property. Connect Person to Project with CONTRIBUTED_TO, and Project to Skill with BUILT_WITH.

Attach embedding vectors to the entities where semantic nuance lives. Rather than storing vectors only on the Person node, attach accomplishment vectors to WORKED_AT edges or dedicated Experience nodes. That separates what someone did at a specific company from who they are as a whole profile.

Keep time ranges explicit on relationship edges. When an engineer moves between firms, the sequence matters for queries like "engineers who joined pre-Series B and stayed through IPO." For the trade-offs between schema topologies, see our guide on vector database vs graph database architectures.

Step 5: Create the vector and BM25 indexes in HelixDB

Running Neo4j for graph queries and Pinecone for vector retrieval means writing custom synchronization layers and managing distributed transactions across two networks. HelixDB puts graph traversal, approximate vector search, and BM25 full-text indexing in a single Rust engine, so all properties and relationships live in one place.

In HelixDB, vector values normalize to float32. Vector properties must be defined at the top level of the node or edge schema. You can set the distance metric to cosine, euclidean, or manhattan depending on your embedding model.

Configure your indexes with the HelixDB TypeScript SDK (@helix-db/helix-db). Here is how you initialize indexes for a talent graph:

import {
  g,
  writeBatch,
  VectorDistanceMetric
} from "@helix-db/helix-db";

// Create nodes with top-level vector properties
const setupQuery = writeBatch([
  // BM25 full-text index on Person name and Company name
  g.createIndex({ label: "Person", property: "name", type: "fulltext" }),
  g.createIndex({ label: "Company", property: "name", type: "fulltext" }),
  
  // Vector index on Experience nodes with 1536 dimensions
  g.createVectorIndex({
    label: "Experience",
    property: "embedding",
    dimensions: 1536,
    metric: VectorDistanceMetric.Cosine
  })
]);

HelixDB runs atomic transactions, so writes to graph relationships, vector indexes, and BM25 indexes commit together. You won't end up with orphaned vector IDs pointing at deleted person records.

Step 6: Prefilter by graph relationship, then vector search (not the other way around)

The standard approach to hybrid search pulls the top 100 nearest vector neighbors first, then discards candidates who fail the relational constraints. Post-filtering like this fails consistently. Say your pool holds 50,000 engineers and only 20 worked at your target companies. A global top-100 vector search may contain zero of them, and your query returns nothing even though qualified candidates exist.

The correct sequence is graph pre-filtering. Isolate candidate nodes through relational traversal first, then run vector search over that subset. When your database scopes vector similarity to entities found during a traversal, it never wastes similarity calculations on irrelevant nodes.

HelixDB natively supports vector pre-filtering scoped by graph relationships. You can traverse from a Company node across WORKED_AT edges to collect candidate Person nodes, then run vector similarity only over the connected Experience nodes. The mechanics are covered in our guide on vector search on graph edges.

Keep HelixDB limits in mind when designing queries. When you run traversal-scoped vector search, HelixDB applies that 800-result ceiling after candidate intersection, and it rejects requests that exceed limits instead of silently clamping. Graph pre-filtering keeps candidate counts small, well inside engine performance bounds.

Step 7: Combine exact-name BM25 lookup with semantic search and test with constraint queries

Production people search has to handle mixed intent. A recruiter might submit "Sarah Chen distributed systems DynamoDB," which needs an exact name match alongside fuzzy semantic criteria. Vector embeddings handle fuzzy concepts well but fail reliably on exact string matching for proper nouns.

Pair HelixDB BM25 text search for proper nouns with vector search for descriptive skills. Large professional networks run similar multi-pass scoring pipelines in production. Use BM25 to pin candidates by exact company, school, or person names, then use vector distance to rank their technical accomplishments.

Test your implementation with hard constraint queries that break pure vector systems:

  1. "Engineers who worked at Snowflake or Databricks and built query optimization engines."

  2. "Founders who previously worked together at Stripe before 2021."

  3. "Kernel engineers who contributed to Linux and know Rust."

If query 1 returns people who never worked at Snowflake, your pre-filtering is misconfigured. If hybrid scoring drops relevant candidates unexpectedly, read our analysis on why does hybrid search return fewer results than you asked for to check candidate deduplication and threshold math.

What to do next: troubleshooting lookalike results and stale edges

Once your people search knowledge graph runs in production, two bugs will dominate your issue tracker: lookalike candidates and stale employment edges.

Lookalike results happen when text chunks contain role titles but no concrete technical detail. If five junior engineers list their title as "Lead Architect" at small local consultancies, embedding models pull them next to genuine principal architects at hyperscalers. Fix this by penalizing sparse chunks at ingest time. If an experience block has fewer than 20 words describing technical output, route it to pure keyword matching and skip the vector property.

Stale edges happen when engineers switch jobs. Without end dates, graphs pile up contradictory records. Store start_date and end_date as integer UNIX timestamps on WORKED_AT edges. HelixDB provides efficient time-range indexing, so your queries can filter out past employment ranges without scanning full node collections.

Audit your entity resolution edge weights monthly. If multiple candidates share common names, require a manual verification flag before merging. Clean boundaries between people, companies, and roles keep the graph trustworthy.

Conclusion

Duct-taping a separate vector database to a graph database creates latency spikes, sync failures, and empty search results. A people search knowledge graph needs unified storage, where graph relationships, vector similarity, and BM25 text indexes all sit inside the same transaction boundary.

HelixDB gives you that in a fast, open-source Rust engine. Stop dropping candidate constraints with post-filtering, and stop running multiple database clusters for one search feature. Clone the open-source repo at github.com/HelixDB/helix-db or build your talent graph directly with the HelixDB TypeScript SDK.

Build with HelixDB

Give your coding agent the setup prompt, or sign up and deploy a database.

Sign up