Be Wellness
Graph RAG - Knowledge Engineer
Skills
Job description
About Chryselys Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights. Role Summary Design, build and operate the Python/FastAPI services that extract entities and relationships from unstructured documents, resolve them to canonical identifiers, maintain the knowledge graph, and serve graph-augmented retrieval alongside vector search for multi-hop and relational questions. Responsibilities Build entity and relation extraction services over unstructured documents — molecules, brands, indications, therapeutic areas, endpoints, claims. Build entity resolution: alias handling, blocking and candidate generation, fuzzy and embedding matching, calibrated thresholds, human review routing. Design and maintain the graph schema and ontology; incremental ingest, node and edge deduplication and merging, provenance on every edge. Fuse graph and vector results into a single ranked, cited context for the retrieval service. Instrument, monitor and support the services in production. Qualifications 5–9 years software engineering, with demonstrable knowledge-graph construction and applied NLP delivered to production. Has built a knowledge graph from unstructured text — not queried an existing one, and not a CRUD application on a graph database. Graph at production scale. Millions of nodes and edges; incremental updates with stable node identity; supernode and traversal-explosion handling with bounded depth and timeouts. Entity resolution at corpus scale. Blocking and candidate generation that avoid O(n²) comparison, with measured precision on a labelled sample. Graph database in production. Neo4j, Amazon Neptune or equivalent; Cypher / openCypher fluency. Ontology and taxonomy modelling. Schema evolution without breaking downstream consumers; judgement on node vs. edge vs. property. Extraction. NER and relation extraction — LLM-based, model-based (spaCy, scispaCy, transformers) or hybrid, with the judgement to choose. Graph vs. vector judgement. Knows where graph retrieval wins — multi-hop, relational, comparative and aggregate questions — and that hybrid is the production norm. Python and FastAPI in production. Python 3.11+, async, Pydantic, Docker, pytest, Git and CI; AWS as a consumer (S3, ECS/EKS, Bedrock, Neptune or self-hosted Neo4j). Preferred Biomedical ontologies and registries: UMLS, MeSH, SNOMED, RxNorm, ICD-10, DrugBank, ChEMBL. Life sciences or pharma domain experience; RDF/SPARQL alongside property graphs. GraphRAG approaches: community detection for corpus-level summarisation, local vs. global search.