PlotDirector/docs/location-significance.md

4.6 KiB
Raw Blame History

Location significance

Audit: the active location import service reads stored scene intelligence, resolves generic references with POV/setting context, normalises aliases, and groups identities using parent-aware keys. Candidate appearances are separate evidence; existing canonical Locations and aliases are consulted before creating proposals. Confidence is the maximum identification confidence. Candidate order was distinct-scene frequency descending, then name. Persisted review candidates retain that order and pending counts drive Review Centre's misleading attention label.

Implementation: keep resolution, merge keys and import decisions unchanged. Compute an explainable, versioned significance assessment from the existing parsed scene results when preparing candidates. Only the actual scene setting inherits scene-wide metrics, plot, knowledge and relationship signals; other places need location-specific consequential evidence. A peak consequential scene outweighs recurrence, which is capped. Store the assessment in the existing candidate JSON, not new location identity fields. Backfill existing candidates deterministically from saved analysis without AI or resetting decisions. Review GETs read persisted assessments. Actual child-setting evidence rolls up to detected parents (for example a house hosting an event in a room), deduplicated by scene; unrelated mentions do not. Hierarchical ordering ranks sibling subtrees while keeping children with their parents; filtered views include ancestors for context. Review Centre counts identified/significant/uncertain separately from pending import choices.

Scoring (version 1): use peak scene evidence, not a sum of mentions. Explicit consequential local evidence or scene purpose contributes 50. Exposition verbs such as “reveals” and “discovers” require consequential subject matter or corroboration from a major plot development (40); they do not automatically establish a major event. A major plot development contributes at least 25 even without event vocabulary. Correlated context signals are capped at 36: an actual setting contributes 12, high relevant metrics 10, plot development 4–6, knowledge changes 3 or 12 for consequential knowledge, consequential relationship evidence 10, and POV association 4. Distinctive named identity contributes 4. Cross-chapter recurrence adds at most 8. Scores cap at 100; 40 marks significant. A significant setting used in six scenes across three chapters becomes Primary Setting. Otherwise three scenes marks Recurring; remaining detections are Incidental. Negated/hypothetical event vocabulary does not establish an event. Scene counts are deduplicated. This conservative heuristic reuses structured evidence; it is not a new semantic AI judgement. It is scoped to the analysed book and does not invent cross-book evidence.

Apply PlotLine/Sql/192_LocationNarrativeSignificance.sql before deploying the application. It adds read/backfill procedures and extends Review Centre's summary; it changes no identity tables. Existing imports can be assessed with Tools/LocationSignificanceBackfill using PLOT_DIRECTOR_SQL_CONNECTION supplied securely:

dotnet run --project Tools/LocationSignificanceBackfill -- <bookId> <userId> <expectedDatabaseName>
dotnet run --project Tools/LocationSignificanceBackfill -- <bookId> <userId> <expectedDatabaseName> --apply

The first command is a dry run. The tool uses the same committed runs (or attached completed analysis fallback) as book resume, verifies database/access and every stable key, and stops before writing if any identity cannot be matched. The apply command modifies assessment JSON only, retaining generation IDs, decisions, merge targets and canonical links. It is repeatable after scoring changes. New candidate generations receive assessments during normal preparation. Missing legacy assessments display “Not yet assessed” and remain available under All detected locations until backfilled. Existing canonical matches still auto-resolve through the established path; significance never changes matching confidence or merge decisions.

Verification on the existing 308-location development import: all stable keys matched; 2 Primary Settings, 54 Significant Locations, 22 Recurring Locations and 230 Incidental Locations. Ten candidates have ambiguous identity/category/parent evidence. All 308 assessments were backfilled with the identity/decision checksum unchanged. The regression harness passes 353 tests, including contextual naming, consequential versus frequent locations, ordinary exposition, parent grouping, uncertainty, pagination and existing merge behaviour.