milindkamat0507 commited on
Commit
1b7aa04
Β·
verified Β·
1 Parent(s): 831088b

Upload tools.py

Browse files
Files changed (1) hide show
  1. tools.py +1203 -0
tools.py ADDED
@@ -0,0 +1,1203 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ tools.py β€” 10 @tool functions for Braun & Clarke (2006) computational
3
+ thematic analysis.
4
+
5
+ Pipeline (called in this order by the LLM agent):
6
+
7
+ 1. load_scopus_csv β€” ingest CSV, strip boilerplate, save .parquet
8
+ 2. run_bertopic_discovery β€” embed β†’ cosine agglomerative cluster (min 3
9
+ members) β†’ centroids β†’ orphan report β†’ 4 charts
10
+ 3. label_topics_with_llm β€” Mistral labels top 100 clusters
11
+ 4. reassign_sentences β€” move orphan/misplaced sentences between clusters
12
+ 5. consolidate_into_themes β€” merge reviewer-approved groups
13
+ 6. compute_saturation β€” coverage %, coherence, balance per theme
14
+ 7. generate_theme_profiles β€” top 5 nearest sentences per theme centroid
15
+ 8. compare_with_taxonomy β€” map themes to PAJAIS 25 categories
16
+ 9. generate_comparison_csv β€” abstract vs title side-by-side
17
+ 10. export_narrative β€” 500-word Section 7 via Mistral
18
+
19
+ Design rules:
20
+
21
+ Every number, percentage, score, or list of sentences presented to the
22
+ reviewer MUST come from a tool β€” never from the LLM's imagination.
23
+
24
+ Deterministic tools (1,2,4,5,6,7,9): same input β†’ same output, every run.
25
+ LLM-dependent tools (3,8,10): grounded in real data passed via prompt,
26
+ but labels/mappings/narrative may vary slightly between runs.
27
+ All LLM-dependent outputs require reviewer approval before advancing.
28
+
29
+ ZERO if/elif/else β€” all decisions by the LLM
30
+ ZERO for/while β€” list(map(...)) and numpy vectorised ops
31
+ ZERO try/except β€” errors surface to the LLM via ToolNode
32
+
33
+ Constants reference:
34
+
35
+ EMBED_MODEL = "all-MiniLM-L6-v2"
36
+ 384d sentence embeddings. Runs locally, no API calls.
37
+ normalize_embeddings=True β†’ cosine similarity = dot product.
38
+
39
+ CLUSTER_THRESHOLD = 0.50
40
+ Cosine distance threshold for Agglomerative Clustering.
41
+ Two sentences must have cosine similarity >= 0.50 to share a code.
42
+ Follows the BERTopic Agglomerative Clustering configuration
43
+ (Grootendorst, 2022) with distance_threshold=0.5 as documented
44
+ in the BERTopic framework. Operationalises Braun & Clarke (2006)
45
+ Phase 2 'Generating Initial Codes' as a reproducible computation.
46
+
47
+ Tighter (e.g. 0.40) β†’ more, finer codes (closer to B&C ideal)
48
+ Looser (e.g. 0.60) β†’ fewer, broader codes
49
+ At 0.50 β€” balanced granularity following BERTopic docs example.
50
+
51
+ MIN_CLUSTER_SIZE = 3
52
+ Clusters with fewer than 3 members are dissolved. Their sentences
53
+ become orphans (label=-1) reported to the reviewer for reassignment.
54
+
55
+ N_CENTROIDS = 200
56
+ Maximum number of clusters saved to summaries.json (and therefore
57
+ labelled and shown in the review table). Set high enough to capture
58
+ all clusters in typical Scopus datasets (1k-5k papers).
59
+ Top clusters extracted for initial discovery report and charts.
60
+
61
+ TOP_TOPICS_LLM = 100
62
+ Maximum clusters sent to Mistral for labelling.
63
+
64
+ NARRATIVE_WORDS = 500
65
+ Target word count for Section 7 narrative.
66
+
67
+ PAJAIS_25
68
+ 25 IS research categories from Jiang et al. (2019).
69
+ Used in Phase 5.5 for taxonomy alignment.
70
+
71
+ BOILERPLATE_PATTERNS (9 regexes)
72
+ Strip publisher noise: copyright, DOI, Elsevier, Springer,
73
+ IEEE, Wiley, Taylor & Francis.
74
+ """
75
+
76
+ from __future__ import annotations
77
+
78
+ import json
79
+ import re
80
+ import numpy as np
81
+ import pandas as pd
82
+ import plotly.graph_objects as go
83
+
84
+ from pathlib import Path
85
+ from langchain_core.tools import tool
86
+ from langchain_mistralai import ChatMistralAI
87
+ from langchain_core.prompts import PromptTemplate
88
+ from langchain_core.output_parsers import JsonOutputParser
89
+ from sentence_transformers import SentenceTransformer
90
+ from sklearn.cluster import AgglomerativeClustering
91
+ from sklearn.metrics.pairwise import cosine_similarity
92
+ from sklearn.preprocessing import normalize
93
+ from sklearn.decomposition import PCA
94
+
95
+
96
+ RUN_CONFIGS = {
97
+ "abstract": ["Abstract"],
98
+ "title": ["Title"],
99
+ }
100
+
101
+ PAJAIS_25 = [
102
+ "Accounting Information Systems",
103
+ "Artificial Intelligence & Expert Systems",
104
+ "Big Data & Analytics",
105
+ "Business Intelligence & Decision Support",
106
+ "Cloud Computing",
107
+ "Cybersecurity & Privacy",
108
+ "Database Management",
109
+ "Digital Transformation",
110
+ "E-Business & E-Commerce",
111
+ "Enterprise Resource Planning",
112
+ "Fintech & Digital Finance",
113
+ "Geographic Information Systems",
114
+ "Health Informatics",
115
+ "Human-Computer Interaction",
116
+ "Information Systems Development",
117
+ "IT Governance & Management",
118
+ "IT Strategy & Competitive Advantage",
119
+ "Knowledge Management",
120
+ "Machine Learning & Deep Learning",
121
+ "Mobile Computing",
122
+ "Natural Language Processing",
123
+ "Recommender Systems",
124
+ "Social Media & Web 2.0",
125
+ "Supply Chain & Logistics IS",
126
+ "Virtual Reality & Augmented Reality",
127
+ ]
128
+
129
+ BOILERPLATE_PATTERNS = [
130
+ r"Β©\s*\d{4}",
131
+ r"all rights reserved",
132
+ r"published by elsevier",
133
+ r"this article is protected",
134
+ r"doi:\s*10\.\d{4,}",
135
+ r"springer nature",
136
+ r"ieee xplore",
137
+ r"wiley online library",
138
+ r"taylor & francis",
139
+ ]
140
+
141
+ BOILERPLATE_RE = re.compile("|".join(BOILERPLATE_PATTERNS), flags=re.IGNORECASE)
142
+ SENTENCE_SPLIT_RE = re.compile(r"(?<=[.!?])\s+")
143
+ EMBED_MODEL = "all-MiniLM-L6-v2"
144
+ N_CENTROIDS = 200
145
+ CLUSTER_THRESHOLD = 0.50
146
+ MIN_CLUSTER_SIZE = 5
147
+ TOP_TOPICS_LLM = 100
148
+ LABEL_BATCH_SIZE = 20
149
+ NARRATIVE_WORDS = 500
150
+
151
+
152
+ def _clean_text(text: str) -> str:
153
+ """Remove publisher boilerplate from a single text string.
154
+
155
+ Applies 9-pattern BOILERPLATE_RE regex to strip copyright notices,
156
+ DOI prefixes, and publisher tags that would pollute embeddings.
157
+
158
+ Args:
159
+ text: Raw abstract or title string.
160
+
161
+ Returns:
162
+ Cleaned string with boilerplate removed and whitespace trimmed.
163
+ """
164
+ return BOILERPLATE_RE.sub("", str(text)).strip()
165
+
166
+
167
+ def _sentence_count(text: str) -> int:
168
+ """Count sentences using regex split on terminal punctuation.
169
+
170
+ Args:
171
+ text: Cleaned abstract or title text.
172
+
173
+ Returns:
174
+ Number of sentences (minimum 1 for any non-empty input).
175
+ """
176
+ return len(SENTENCE_SPLIT_RE.split(text.strip()))
177
+
178
+
179
+ def _embed(texts: list[str]) -> np.ndarray:
180
+ """Embed texts into 384d L2-normalized unit vectors.
181
+
182
+ Uses SentenceTransformer('all-MiniLM-L6-v2') locally β€” no API calls.
183
+ normalize_embeddings=True ensures cosine_similarity = dot product.
184
+
185
+ Args:
186
+ texts: List of N cleaned text strings.
187
+
188
+ Returns:
189
+ np.ndarray shape (N, 384), dtype float32, L2-normalized.
190
+ """
191
+ model = SentenceTransformer(EMBED_MODEL)
192
+ raw = model.encode(texts, show_progress_bar=False, normalize_embeddings=True)
193
+ return np.array(raw, dtype=np.float32)
194
+
195
+
196
+ def _cosine_cluster(matrix: np.ndarray, threshold: float, min_size: int) -> np.ndarray:
197
+ """Cluster embeddings using agglomerative cosine clustering.
198
+
199
+ Works DIRECTLY in 384d space β€” no UMAP. After clustering, any cluster
200
+ with fewer than min_size members is dissolved: its sentences get
201
+ label=-1 (orphan) and are reported to the reviewer for reassignment.
202
+
203
+ Algorithm:
204
+ 1. Start: every text is its own cluster.
205
+ 2. Merge the two closest clusters (average cosine distance).
206
+ 3. Repeat until smallest distance exceeds threshold.
207
+ 4. Post-process: dissolve clusters smaller than min_size.
208
+
209
+ Args:
210
+ matrix: (N, 384) embedding matrix, L2-normalized.
211
+ threshold: Max cosine distance for merging (0.7 β†’ ~100 clusters).
212
+ min_size: Minimum members per cluster (3). Smaller β†’ orphan.
213
+
214
+ Returns:
215
+ np.ndarray shape (N,) with integer labels. -1 = orphan.
216
+ """
217
+ normed = normalize(matrix, norm="l2")
218
+ model = AgglomerativeClustering(
219
+ n_clusters=None,
220
+ metric="cosine",
221
+ linkage="average",
222
+ distance_threshold=threshold,
223
+ )
224
+ labels = model.fit_predict(normed).astype(int)
225
+ unique, counts = np.unique(labels, return_counts=True)
226
+ small_clusters = unique[counts < min_size]
227
+ return np.where(np.isin(labels, small_clusters), -1, labels)
228
+
229
+
230
+ def _centroid(vecs: np.ndarray) -> np.ndarray:
231
+ """Compute L2-normalized centroid (average direction in 384d space).
232
+
233
+ Args:
234
+ vecs: (M, 384) matrix of member embeddings for one cluster.
235
+
236
+ Returns:
237
+ 1d np.ndarray shape (384,), L2-normalized.
238
+ """
239
+ return normalize(vecs.mean(axis=0, keepdims=True), norm="l2")[0]
240
+
241
+
242
+ def _top_n_centroids(matrix: np.ndarray, labels: np.ndarray, n: int) -> list[dict]:
243
+ """Extract N largest clusters by size and compute their centroids.
244
+
245
+ Excludes orphans (label=-1) from the ranking.
246
+
247
+ Args:
248
+ matrix: (N, 384) full embedding matrix.
249
+ labels: (N,) integer cluster labels (-1 = orphan).
250
+ n: How many top clusters to return.
251
+
252
+ Returns:
253
+ List of N dicts with: label, size, indices, centroid.
254
+ """
255
+ valid_mask = labels >= 0
256
+ valid_labels = labels[valid_mask]
257
+ unique, counts = np.unique(valid_labels, return_counts=True)
258
+ order = np.argsort(counts)[::-1][:n]
259
+ top_labels = unique[order]
260
+
261
+ def _build(lbl: int) -> dict:
262
+ """Build summary dict for one cluster."""
263
+ idx = np.where(labels == lbl)[0].tolist()
264
+ return {
265
+ "label": int(lbl),
266
+ "size": len(idx),
267
+ "indices": idx,
268
+ "centroid": _centroid(matrix[idx]),
269
+ }
270
+
271
+ return list(map(_build, top_labels))
272
+
273
+
274
+ def _mistral_chain(template_str: str):
275
+ """Create PromptTemplate β†’ ChatMistralAI β†’ JsonOutputParser chain.
276
+
277
+ Args:
278
+ template_str: Prompt template with {variable} placeholders.
279
+
280
+ Returns:
281
+ LangChain Runnable chain that accepts dict and returns parsed JSON.
282
+ """
283
+ llm = ChatMistralAI(
284
+ model="mistral-large-latest",
285
+ temperature=0,
286
+ timeout=240,
287
+ max_retries=3,
288
+ )
289
+ prompt = PromptTemplate.from_template(template_str)
290
+ return prompt | llm | JsonOutputParser()
291
+
292
+
293
+ def _dark_layout(title: str) -> dict:
294
+ """Return Plotly layout dict with dark theme styling.
295
+
296
+ Args:
297
+ title: Chart title string.
298
+
299
+ Returns:
300
+ Dict for fig.update_layout(**_dark_layout("...")).
301
+ """
302
+ return dict(
303
+ title=title, paper_bgcolor="#0F172A", plot_bgcolor="#0F172A",
304
+ font=dict(color="#CBD5E1", family="Sora,sans-serif"),
305
+ margin=dict(t=50, b=40, l=40, r=20),
306
+ )
307
+
308
+
309
+ @tool
310
+ def load_scopus_csv(csv_path: str, run_mode: str = "abstract") -> str:
311
+ """Load a Scopus CSV, count papers/sentences, apply boilerplate filter.
312
+
313
+ Phase 1 β€” Familiarisation with the Data. DETERMINISTIC.
314
+
315
+ Steps:
316
+ 1. Read CSV, drop rows where target column is null
317
+ 2. Apply 9-pattern boilerplate regex to clean each text
318
+ 3. Count sentences per paper
319
+ 4. Save cleaned DataFrame as .parquet
320
+
321
+ Args:
322
+ csv_path: Path to raw Scopus CSV.
323
+ run_mode: 'abstract' or 'title'.
324
+
325
+ Returns:
326
+ JSON: total_papers, total_sentences, columns_used,
327
+ boilerplate_removed, cleaned_parquet, run_mode.
328
+ """
329
+ cols = RUN_CONFIGS[run_mode]
330
+ target = cols[0]
331
+
332
+ df = pd.read_csv(csv_path).dropna(subset=[target]).reset_index(drop=True)
333
+ raw_texts = df[target].tolist()
334
+ cleaned_texts = list(map(_clean_text, raw_texts))
335
+
336
+ boilerplate_removed = sum(map(
337
+ lambda pair: int(pair[0] != pair[1]),
338
+ zip(raw_texts, cleaned_texts),
339
+ ))
340
+
341
+ df[f"{target}_clean"] = cleaned_texts
342
+ df["sentence_count"] = list(map(_sentence_count, cleaned_texts))
343
+
344
+ out_path = Path(csv_path).with_suffix(".clean.parquet")
345
+ df.to_parquet(out_path, index=False)
346
+
347
+ return json.dumps({
348
+ "total_papers": len(df),
349
+ "total_sentences": int(df["sentence_count"].sum()),
350
+ "columns_used": cols,
351
+ "boilerplate_removed": boilerplate_removed,
352
+ "cleaned_parquet": str(out_path),
353
+ "run_mode": run_mode,
354
+ }, indent=2)
355
+
356
+
357
+ @tool
358
+ def run_bertopic_discovery(parquet_path: str, run_mode: str = "abstract") -> str:
359
+ """Embed texts, cluster them, report orphans, generate charts.
360
+
361
+ Phase 2 β€” Generating Initial Codes. DETERMINISTIC.
362
+
363
+ Steps:
364
+ 1. Load cleaned parquet, drop Author Keywords columns (RULE 8)
365
+ 2. Embed all texts β†’ N x 384 matrix of unit vectors
366
+ 3. Save embedding matrix as .emb.npy
367
+ 4. Cluster in 384d space (NO UMAP), min 3 members per cluster
368
+ 5. Sentences in clusters < 3 members become orphans (label=-1)
369
+ 6. Extract top-N clusters by size, compute centroids
370
+ 7. Save summaries.json with clusters + orphan list
371
+ 8. Generate 4 Plotly HTML charts
372
+
373
+ Args:
374
+ parquet_path: Path to .clean.parquet from load_scopus_csv.
375
+ run_mode: 'abstract' or 'title'.
376
+
377
+ Returns:
378
+ JSON: total_clusters, orphan_count, summaries_json, embeddings_npy,
379
+ charts dict.
380
+ """
381
+ cols = RUN_CONFIGS[run_mode]
382
+ target = f"{cols[0]}_clean"
383
+
384
+ df = pd.read_parquet(parquet_path).drop(
385
+ columns=[c for c in pd.read_parquet(parquet_path).columns
386
+ if re.search(r"keyword|author", c, re.I)],
387
+ errors="ignore",
388
+ )
389
+
390
+ paper_texts = df[target].tolist()
391
+
392
+ sentence_records = list(filter(
393
+ lambda r: len(r["text"].split()) >= 5,
394
+ [
395
+ {"paper_idx": paper_i, "sent_idx": sent_i, "text": sent.strip()}
396
+ for paper_i, paper_text in enumerate(paper_texts)
397
+ for sent_i, sent in enumerate(SENTENCE_SPLIT_RE.split(paper_text or ""))
398
+ if sent.strip()
399
+ ],
400
+ ))
401
+
402
+ texts = list(map(lambda r: r["text"], sentence_records))
403
+ paper_idx = list(map(lambda r: r["paper_idx"], sentence_records))
404
+ embeddings = _embed(texts)
405
+ base = Path(parquet_path).parent
406
+
407
+ np.save(str(base / Path(parquet_path).stem) + ".emb.npy", embeddings)
408
+ (base / "sentences.json").write_text(json.dumps({
409
+ "texts": texts,
410
+ "paper_idx": paper_idx,
411
+ }))
412
+
413
+ labels = _cosine_cluster(embeddings, CLUSTER_THRESHOLD, MIN_CLUSTER_SIZE)
414
+ orphan_idx = np.where(labels == -1)[0].tolist()
415
+ orphan_count = len(orphan_idx)
416
+ valid_count = int((labels >= 0).sum())
417
+ n_clusters = int(np.unique(labels[labels >= 0]).shape[0])
418
+ n_papers = len(set(paper_idx))
419
+ n_sentences = len(texts)
420
+ top_centroids = _top_n_centroids(embeddings, labels, N_CENTROIDS)
421
+
422
+ def _topic_row(tc: dict) -> dict:
423
+ """Convert centroid dict into summary row for summaries.json."""
424
+ return {
425
+ "topic_id": tc["label"],
426
+ "size": tc["size"],
427
+ "representative": texts[tc["indices"][0]][:200],
428
+ "indices": tc["indices"],
429
+ }
430
+
431
+ summaries = list(map(_topic_row, top_centroids))
432
+
433
+ orphans = list(map(
434
+ lambda i: {"sentence_idx": int(i), "text": texts[i][:200]},
435
+ orphan_idx,
436
+ ))
437
+
438
+ output = {"clusters": summaries, "orphans": orphans}
439
+ (base / "summaries.json").write_text(json.dumps(output, indent=2))
440
+
441
+ unique, counts = np.unique(labels[labels >= 0], return_counts=True)
442
+ order = np.argsort(counts)[::-1][:20]
443
+ c1 = go.Figure(go.Bar(
444
+ x=list(map(str, unique[order])), y=counts[order].tolist(),
445
+ marker_color="#3B82F6", text=counts[order].tolist(), textposition="outside",
446
+ ))
447
+ c1.update_layout(**_dark_layout("Topic Size Distribution (Top 20)"),
448
+ xaxis=dict(showgrid=False),
449
+ yaxis=dict(showgrid=True, gridcolor="#1E293B"))
450
+ c1.write_html(str(base / "chart_topic_sizes.html"))
451
+
452
+ centroid_matrix = np.vstack([tc["centroid"] for tc in top_centroids])
453
+ sim_matrix = cosine_similarity(centroid_matrix)
454
+ clabels = list(map(lambda tc: f"T{tc['label']}", top_centroids))
455
+ c2 = go.Figure(go.Heatmap(z=sim_matrix, x=clabels, y=clabels, colorscale="Blues"))
456
+ c2.update_layout(**_dark_layout("Top-5 Centroid Cosine Similarity"))
457
+ c2.write_html(str(base / "chart_centroid_heatmap.html"))
458
+
459
+ sc = df.get("sentence_count", pd.Series([0] * len(df))).tolist()
460
+ c3 = go.Figure(go.Histogram(x=sc, nbinsx=40, marker_color="#22D3EE"))
461
+ c3.update_layout(**_dark_layout("Sentence Count Distribution"),
462
+ xaxis=dict(showgrid=False),
463
+ yaxis=dict(showgrid=True, gridcolor="#1E293B"))
464
+ c3.write_html(str(base / "chart_sentence_distribution.html"))
465
+
466
+ coords = PCA(n_components=2).fit_transform(centroid_matrix)
467
+ point_text = list(map(lambda tc: f"T{tc['label']}({tc['size']})", top_centroids))
468
+ c4 = go.Figure(go.Scatter(
469
+ x=coords[:, 0].tolist(), y=coords[:, 1].tolist(),
470
+ mode="markers+text", text=point_text, textposition="top center",
471
+ marker=dict(size=12, color="#F59E0B", line=dict(width=1, color="#0F172A")),
472
+ ))
473
+ c4.update_layout(**_dark_layout("Top-5 Centroids β€” PCA Projection"))
474
+ c4.write_html(str(base / "chart_centroid_pca.html"))
475
+
476
+ emb_path = str(base / Path(parquet_path).stem) + ".emb.npy"
477
+ return json.dumps({
478
+ "total_clusters": n_clusters,
479
+ "orphan_count": orphan_count,
480
+ "valid_sentences": valid_count,
481
+ "total_sentences": n_sentences,
482
+ "total_papers": n_papers,
483
+ "top_centroids": N_CENTROIDS,
484
+ "summaries_json": str(base / "summaries.json"),
485
+ "embeddings_npy": emb_path,
486
+ "needs_review": True,
487
+ "charts": {
488
+ "topic_sizes": str(base / "chart_topic_sizes.html"),
489
+ "centroid_heatmap": str(base / "chart_centroid_heatmap.html"),
490
+ "sentence_dist": str(base / "chart_sentence_distribution.html"),
491
+ "centroid_pca": str(base / "chart_centroid_pca.html"),
492
+ },
493
+ }, indent=2)
494
+
495
+
496
+ @tool
497
+ def label_topics_with_llm(summaries_json_path: str) -> str:
498
+ """Send top-100 topic summaries to Mistral for labelling.
499
+
500
+ Phase 2 β€” Naming Initial Codes. LLM-DEPENDENT (grounded in real data extracts).
501
+ NOTE: Prefer run_phase_1_and_2 for the standard Phase 2 entry point.
502
+ This tool is kept for backwards compatibility and edge-case re-labelling.
503
+
504
+ Args:
505
+ summaries_json_path: Path to summaries.json.
506
+
507
+ Returns:
508
+ JSON: labelled_topics count + output path. needs_review=True.
509
+ """
510
+ data = json.loads(Path(summaries_json_path).read_text())
511
+ summaries = data.get("clusters", data)[:TOP_TOPICS_LLM]
512
+ result = _label_summaries_with_mistral(summaries)
513
+ out_path = Path(summaries_json_path).parent / "topic_labels.json"
514
+ out_path.write_text(json.dumps(result, indent=2))
515
+
516
+ return json.dumps({
517
+ "labelled_topics": len(result),
518
+ "output": str(out_path),
519
+ "needs_review": True,
520
+ }, indent=2)
521
+
522
+
523
+ def _label_summaries_with_mistral(summaries: list[dict]) -> list[dict]:
524
+ """Internal helper: send cluster summaries to Mistral for labelling in batches.
525
+
526
+ Batches into groups of LABEL_BATCH_SIZE (20) to avoid Mistral API
527
+ timeouts that occur when sending all 100 summaries in one prompt.
528
+ Each batch is a separate API call; results are concatenated.
529
+
530
+ Returns a list of dicts with topic_id, label, rationale, confidence.
531
+ Used by both label_topics_with_llm and run_phase_1_and_2.
532
+ """
533
+ template = (
534
+ "You are a scientific topic labelling expert.\n\n"
535
+ "Below are {n} topic summaries from a BERTopic analysis of academic papers.\n"
536
+ "Each summary has: topic_id, size, representative text.\n\n"
537
+ "{summaries}\n\n"
538
+ "For EACH topic return a JSON array where every element has:\n"
539
+ " topic_id : integer (copy from input)\n"
540
+ " label : 2-5 word snake_case topic label\n"
541
+ " rationale : one sentence justification\n"
542
+ " confidence : float 0.0-1.0\n\n"
543
+ "Return ONLY the JSON array β€” no markdown, no preamble."
544
+ )
545
+ chain = _mistral_chain(template)
546
+ batches = [summaries[i:i + LABEL_BATCH_SIZE]
547
+ for i in range(0, len(summaries), LABEL_BATCH_SIZE)]
548
+ results = list(map(
549
+ lambda batch: chain.invoke({
550
+ "n": len(batch),
551
+ "summaries": json.dumps(batch, indent=2),
552
+ }),
553
+ batches,
554
+ ))
555
+ return sum(results, [])
556
+
557
+
558
+ @tool
559
+ def run_phase_1_and_2(csv_path: str, run_mode: str = "abstract") -> str:
560
+ """Execute Phase 1 (Familiarisation) + Phase 2 (Generating Initial Codes)
561
+ in a SINGLE tool call. The canonical entry point for analysis.
562
+
563
+ This is the ONE tool the agent should call when the user clicks
564
+ "Run analysis on abstracts" or "Run analysis on titles".
565
+
566
+ Internally performs:
567
+ 1. Phase 1 β€” Familiarisation with the Data:
568
+ - Load Scopus CSV, drop rows with empty target column
569
+ - Apply boilerplate regex cleaner
570
+ - Save .clean.parquet
571
+
572
+ 2. Phase 2a β€” Sentence Splitting & Embedding:
573
+ - Split each cleaned data item into sentences
574
+ - Filter to sentences with >= 5 words
575
+ - Embed with Sentence-BERT all-MiniLM-L6-v2
576
+ - Save .emb.npy + sentences.json
577
+
578
+ 3. Phase 2b β€” Cosine Agglomerative Clustering:
579
+ - sklearn.cluster.AgglomerativeClustering with metric='cosine',
580
+ linkage='average', distance_threshold=0.50
581
+ - Enforce minimum 5 extracts per code (smaller β†’ orphan)
582
+ - Save summaries.json (top N centroids)
583
+
584
+ 4. Phase 2c β€” LLM Naming via Mistral:
585
+ - Top 100 codes (by size) sent to Mistral for snake_case labels
586
+ - Save topic_labels.json
587
+
588
+ All checkpoint files are saved to the SAME directory as csv_path,
589
+ forming a workspace that downstream tools can discover via workspace_dir.
590
+
591
+ Args:
592
+ csv_path: Path to raw Scopus CSV.
593
+ run_mode: 'abstract' or 'title' β€” which column to analyse.
594
+
595
+ Returns:
596
+ JSON with combined Phase 1 + Phase 2 metrics:
597
+ phase_1: data_items, data_extracts, boilerplate_removed
598
+ phase_2: initial_codes, orphan_extracts, labelled_count
599
+ workspace_dir: directory containing all checkpoints
600
+ needs_review: True (Phase 2 STOP gate awaits)
601
+ """
602
+ cols = RUN_CONFIGS[run_mode]
603
+ target = cols[0]
604
+
605
+ df = pd.read_csv(csv_path).dropna(subset=[target]).reset_index(drop=True)
606
+ raw_texts = df[target].tolist()
607
+ cleaned_texts = list(map(_clean_text, raw_texts))
608
+
609
+ boilerplate_removed = sum(map(
610
+ lambda pair: int(pair[0] != pair[1]),
611
+ zip(raw_texts, cleaned_texts),
612
+ ))
613
+
614
+ df[f"{target}_clean"] = cleaned_texts
615
+ df["sentence_count"] = list(map(_sentence_count, cleaned_texts))
616
+
617
+ workspace = Path(csv_path).parent
618
+ parquet_path = workspace / (Path(csv_path).stem + ".clean.parquet")
619
+ df.to_parquet(parquet_path, index=False)
620
+
621
+ sentence_records = list(filter(
622
+ lambda r: len(r["text"].split()) >= 5,
623
+ [
624
+ {"paper_idx": paper_i, "sent_idx": sent_i, "text": sent.strip()}
625
+ for paper_i, paper_text in enumerate(cleaned_texts)
626
+ for sent_i, sent in enumerate(SENTENCE_SPLIT_RE.split(paper_text or ""))
627
+ if sent.strip()
628
+ ],
629
+ ))
630
+
631
+ texts = list(map(lambda r: r["text"], sentence_records))
632
+ paper_idx = list(map(lambda r: r["paper_idx"], sentence_records))
633
+ embeddings = _embed(texts)
634
+
635
+ np.save(str(workspace / Path(csv_path).stem) + ".emb.npy", embeddings)
636
+ (workspace / "sentences.json").write_text(json.dumps({
637
+ "texts": texts,
638
+ "paper_idx": paper_idx,
639
+ }))
640
+
641
+ labels = _cosine_cluster(embeddings, CLUSTER_THRESHOLD, MIN_CLUSTER_SIZE)
642
+ orphan_idx = np.where(labels == -1)[0].tolist()
643
+ orphan_count = len(orphan_idx)
644
+ valid_count = int((labels >= 0).sum())
645
+ n_clusters = int(np.unique(labels[labels >= 0]).shape[0])
646
+ top_centroids = _top_n_centroids(embeddings, labels, N_CENTROIDS)
647
+
648
+ summaries = list(map(
649
+ lambda tc: {
650
+ "topic_id": int(tc["label"]),
651
+ "size": tc["size"],
652
+ "representative": texts[tc["indices"][0]][:200],
653
+ "indices": tc["indices"],
654
+ },
655
+ top_centroids,
656
+ ))
657
+ orphans = list(map(
658
+ lambda i: {"sentence_idx": int(i), "text": texts[i][:200]},
659
+ orphan_idx,
660
+ ))
661
+ (workspace / "summaries.json").write_text(json.dumps({
662
+ "clusters": summaries,
663
+ "orphans": orphans,
664
+ }, indent=2))
665
+
666
+ labelling_input = list(map(
667
+ lambda s: {k: v for k, v in s.items() if k != "indices"},
668
+ summaries[:TOP_TOPICS_LLM],
669
+ ))
670
+ labelled = _label_summaries_with_mistral(labelling_input)
671
+
672
+ indices_by_id = {s["topic_id"]: s["indices"] for s in summaries}
673
+ enriched = list(map(
674
+ lambda l: {**l,
675
+ "topic_id": int(l.get("topic_id", -1)),
676
+ "size": len(indices_by_id.get(int(l.get("topic_id", -1)), [])),
677
+ "indices": indices_by_id.get(int(l.get("topic_id", -1)), [])},
678
+ labelled,
679
+ ))
680
+ (workspace / "topic_labels.json").write_text(json.dumps(enriched, indent=2))
681
+
682
+ return json.dumps({
683
+ "phase_1": {
684
+ "data_items": len(df),
685
+ "data_extracts": len(texts),
686
+ "boilerplate_removed": boilerplate_removed,
687
+ },
688
+ "phase_2": {
689
+ "initial_codes": n_clusters,
690
+ "labelled_count": len(enriched),
691
+ "orphan_extracts": orphan_count,
692
+ "min_cluster": MIN_CLUSTER_SIZE,
693
+ },
694
+ "workspace_dir": str(workspace),
695
+ "summaries_json": str(workspace / "summaries.json"),
696
+ "labels_json": str(workspace / "topic_labels.json"),
697
+ "embeddings_npy": str(workspace / Path(csv_path).stem) + ".emb.npy",
698
+ "sentences_json": str(workspace / "sentences.json"),
699
+ "needs_review": True,
700
+ }, indent=2)
701
+
702
+
703
+ @tool
704
+ def reassign_sentences(
705
+ summaries_json_path: str,
706
+ embeddings_npy_path: str,
707
+ move_instructions: list[dict],
708
+ ) -> str:
709
+ """Move orphan or misplaced sentences between clusters.
710
+
711
+ Phase 2 β€” Reassigning orphan data extracts. DETERMINISTIC.
712
+
713
+ The reviewer specifies moves as a list of dicts:
714
+ [{"sentence_idx": 42, "to_cluster": 3},
715
+ {"sentence_idx": 99, "to_cluster": "new"}]
716
+
717
+ For "new" targets, a fresh cluster ID is assigned.
718
+ After all moves, centroids are recomputed for affected clusters.
719
+
720
+ Steps:
721
+ 1. Load summaries.json and embeddings
722
+ 2. Apply move instructions
723
+ 3. Update cluster assignments
724
+ 4. Recompute centroids for affected clusters
725
+ 5. Save updated summaries.json
726
+
727
+ Args:
728
+ summaries_json_path: Path to summaries.json.
729
+ embeddings_npy_path: Path to .emb.npy.
730
+ move_instructions: List of dicts with sentence_idx (int) and
731
+ to_cluster (int or "new") keys.
732
+
733
+ Returns:
734
+ JSON: moves_applied count, orphans_remaining, updated summaries path.
735
+ """
736
+ data = json.loads(Path(summaries_json_path).read_text())
737
+ embeddings = np.load(embeddings_npy_path)
738
+ moves = move_instructions
739
+ clusters = data.get("clusters", [])
740
+ orphans = data.get("orphans", [])
741
+
742
+ all_indices = {}
743
+ list(map(
744
+ lambda c: all_indices.update({idx: c["topic_id"] for idx in c.get("indices", [])}),
745
+ clusters,
746
+ ))
747
+
748
+ max_id = max(map(lambda c: c.get("topic_id", 0), clusters), default=0)
749
+ new_id_counter = [max_id + 1]
750
+
751
+ def _apply_move(m: dict) -> dict:
752
+ """Apply one move instruction, return the resolved target cluster ID."""
753
+ s_idx = m["sentence_idx"]
754
+ target = m["to_cluster"]
755
+ resolved = (target == "new") and new_id_counter.__setitem__(0, new_id_counter[0] + 1) or target
756
+ final_id = new_id_counter[0] - 1 * (target == "new") + target * (target != "new")
757
+ all_indices[s_idx] = int(target) * (target != "new") + new_id_counter[0] * (target == "new")
758
+ return {"sentence_idx": s_idx, "assigned_to": all_indices[s_idx]}
759
+
760
+ applied = list(map(_apply_move, moves))
761
+
762
+ unique_clusters = set(all_indices.values())
763
+
764
+ def _rebuild_cluster(cid: int) -> dict:
765
+ """Rebuild a cluster dict from the updated index map."""
766
+ idx = [k for k, v in all_indices.items() if v == cid]
767
+ vecs = embeddings[idx or [0]]
768
+ return {
769
+ "topic_id": int(cid),
770
+ "size": len(idx),
771
+ "representative": "",
772
+ "indices": idx,
773
+ "centroid": _centroid(vecs).tolist(),
774
+ }
775
+
776
+ updated_clusters = list(map(_rebuild_cluster, sorted(unique_clusters)))
777
+ remaining_orphan_idx = [o["sentence_idx"] for o in orphans
778
+ if o["sentence_idx"] not in all_indices]
779
+
780
+ output = {
781
+ "clusters": updated_clusters,
782
+ "orphans": list(map(
783
+ lambda i: {"sentence_idx": i, "text": ""},
784
+ remaining_orphan_idx,
785
+ )),
786
+ }
787
+ Path(summaries_json_path).write_text(json.dumps(output, indent=2))
788
+
789
+ return json.dumps({
790
+ "moves_applied": len(applied),
791
+ "orphans_remaining": len(remaining_orphan_idx),
792
+ "summaries_json": summaries_json_path,
793
+ "needs_review": True,
794
+ }, indent=2)
795
+
796
+
797
+ @tool
798
+ def consolidate_into_themes(
799
+ labels_json_path: str,
800
+ embeddings_npy_path: str,
801
+ approved_topic_ids: list[list[int]],
802
+ ) -> str:
803
+ """Merge approved topic groups into consolidated themes.
804
+
805
+ Phase 3 β€” Searching for Themes. DETERMINISTIC.
806
+
807
+ Steps:
808
+ 1. Load topic_labels.json and embedding matrix
809
+ 2. Pool all member embeddings per group
810
+ 3. Compute fresh L2-normalized centroid per merged group
811
+ 4. Build theme name from joined sub-labels
812
+ 5. Save themes.json
813
+
814
+ Args:
815
+ labels_json_path: Path to topic_labels.json.
816
+ embeddings_npy_path: Path to .emb.npy.
817
+ approved_topic_ids: List of lists of initial-code IDs.
818
+ Each inner list is one candidate theme.
819
+ Example: [[0,1,2],[3,4],[5]] creates 3
820
+ candidate themes from 6 initial codes.
821
+
822
+ Returns:
823
+ JSON: themes_created count + themes_json path. needs_review=True.
824
+ """
825
+ labels_data = json.loads(Path(labels_json_path).read_text())
826
+ embeddings = np.load(embeddings_npy_path)
827
+ groups = approved_topic_ids
828
+ label_map = {item["topic_id"]: item for item in labels_data}
829
+
830
+ def _merge_group(group_ids: list[int]) -> dict:
831
+ """Merge topic IDs into one theme, recompute centroid."""
832
+ members = [m for m in map(label_map.get, group_ids) if m is not None]
833
+ all_idx = sum(map(lambda m: m.get("indices", []), members), [])
834
+ vecs = embeddings[all_idx or [0]]
835
+ centroid = _centroid(vecs)
836
+ sub_labels = list(map(lambda m: m.get("label", ""), members))
837
+ theme_name = "_".join(
838
+ dict.fromkeys(sum(map(lambda lbl: lbl.split("_"), sub_labels), []))
839
+ )[:60]
840
+ return {
841
+ "theme_id": group_ids[0],
842
+ "theme_label": theme_name,
843
+ "merged_ids": group_ids,
844
+ "total_papers": len(set(all_idx)),
845
+ "indices": all_idx,
846
+ "centroid": centroid.tolist(),
847
+ }
848
+
849
+ themes = list(map(_merge_group, groups))
850
+ out_path = Path(labels_json_path).parent / "themes.json"
851
+ out_path.write_text(json.dumps(themes, indent=2))
852
+
853
+ return json.dumps({
854
+ "themes_created": len(themes),
855
+ "themes_json": str(out_path),
856
+ "needs_review": True,
857
+ }, indent=2)
858
+
859
+
860
+ @tool
861
+ def compute_saturation(
862
+ themes_json_path: str,
863
+ embeddings_npy_path: str,
864
+ total_papers: int,
865
+ ) -> str:
866
+ """Compute saturation metrics per theme: coverage, coherence, balance.
867
+
868
+ Phase 4 β€” Reviewing Themes. DETERMINISTIC.
869
+
870
+ Every number in the output is computed by numpy β€” the LLM never
871
+ calculates these values. This eliminates hallucination risk for
872
+ percentages, scores, and ratios.
873
+
874
+ Metrics per theme:
875
+ coverage = papers_in_theme / total_papers (exact percentage)
876
+ coherence = mean pairwise cosine similarity of member embeddings
877
+ (1.0 = all identical, 0.0 = orthogonal)
878
+
879
+ Global metrics:
880
+ total_coverage = papers in at least one theme / total_papers
881
+ balance_ratio = largest_theme / smallest_theme
882
+ mean_coherence = average of per-theme coherence scores
883
+
884
+ Args:
885
+ themes_json_path: Path to themes.json.
886
+ embeddings_npy_path: Path to .emb.npy.
887
+ total_papers: Total papers in corpus (from Phase 1 stats).
888
+
889
+ Returns:
890
+ JSON: per-theme metrics + global metrics. needs_review=True.
891
+ """
892
+ themes = json.loads(Path(themes_json_path).read_text())
893
+ embeddings = np.load(embeddings_npy_path)
894
+
895
+ def _theme_metrics(t: dict) -> dict:
896
+ """Compute coverage and coherence for one theme."""
897
+ idx = t.get("indices", [])
898
+ size = len(idx)
899
+ vecs = embeddings[idx or [0]]
900
+ sim = cosine_similarity(vecs)
901
+ n = len(vecs)
902
+ coherence = float(
903
+ (sim.sum() - n) / max(n * (n - 1), 1)
904
+ )
905
+ return {
906
+ "theme_id": t.get("theme_id", 0),
907
+ "theme_label": t.get("theme_label", ""),
908
+ "papers": size,
909
+ "coverage_pct": round(size / max(total_papers, 1) * 100, 2),
910
+ "coherence": round(coherence, 4),
911
+ }
912
+
913
+ per_theme = list(map(_theme_metrics, themes))
914
+
915
+ all_paper_idx = set(sum(map(lambda t: t.get("indices", []), themes), []))
916
+ sizes = list(map(lambda m: m["papers"], per_theme))
917
+ coherences = list(map(lambda m: m["coherence"], per_theme))
918
+
919
+ global_metrics = {
920
+ "total_coverage_pct": round(len(all_paper_idx) / max(total_papers, 1) * 100, 2),
921
+ "balance_ratio": round(max(sizes, default=1) / max(min(sizes, default=1), 1), 2),
922
+ "mean_coherence": round(sum(coherences) / max(len(coherences), 1), 4),
923
+ "theme_count": len(themes),
924
+ }
925
+
926
+ out_path = Path(themes_json_path).parent / "saturation.json"
927
+ result = {"per_theme": per_theme, "global": global_metrics}
928
+ out_path.write_text(json.dumps(result, indent=2))
929
+
930
+ return json.dumps({
931
+ **global_metrics,
932
+ "per_theme": per_theme,
933
+ "saturation_json": str(out_path),
934
+ "needs_review": True,
935
+ }, indent=2)
936
+
937
+
938
+ @tool
939
+ def generate_theme_profiles(
940
+ themes_json_path: str,
941
+ embeddings_npy_path: str,
942
+ texts_parquet_path: str,
943
+ run_mode: str = "abstract",
944
+ ) -> str:
945
+ """Generate profile cards with top-5 nearest sentences per theme.
946
+
947
+ Phase 5 β€” Defining and Naming Themes. DETERMINISTIC.
948
+
949
+ For each theme centroid, computes cosine similarity against ALL
950
+ embeddings and returns the 5 closest sentences. These are the
951
+ REAL sentences from the corpus β€” not generated, not recalled
952
+ from conversation history. The reviewer uses these to decide
953
+ on final theme names.
954
+
955
+ Steps:
956
+ 1. Load themes.json with centroids
957
+ 2. Load full embedding matrix (sentence-level)
958
+ 3. Load sentences.json (the EXACT sentences that were embedded)
959
+ 4. For each theme: cosine_similarity(centroid, all_embeddings)
960
+ 5. Take top 5 by similarity score
961
+ 6. Return exact sentence text + similarity score
962
+ 7. Save profiles.json
963
+
964
+ Args:
965
+ themes_json_path: Path to themes.json.
966
+ embeddings_npy_path: Path to .emb.npy.
967
+ texts_parquet_path: Path to .clean.parquet (kept for compatibility,
968
+ but sentences are now loaded from sentences.json
969
+ which lives in the same directory).
970
+ run_mode: 'abstract' or 'title'.
971
+
972
+ Returns:
973
+ JSON: profiles list with top-5 sentences per theme. needs_review=True.
974
+ """
975
+ themes = json.loads(Path(themes_json_path).read_text())
976
+ embeddings = np.load(embeddings_npy_path)
977
+ sentences_path = Path(themes_json_path).parent / "sentences.json"
978
+ sentences_data = json.loads(sentences_path.read_text())
979
+ texts = sentences_data["texts"]
980
+
981
+ def _profile(t: dict) -> dict:
982
+ """Build a profile card for one theme: centroid β†’ top 5 nearest."""
983
+ centroid = np.array(t["centroid"]).reshape(1, -1)
984
+ sims = cosine_similarity(centroid, embeddings)[0]
985
+ top5_idx = np.argsort(sims)[::-1][:5].tolist()
986
+ top5 = list(map(
987
+ lambda i: {
988
+ "sentence_idx": i,
989
+ "text": texts[i][:300],
990
+ "similarity": round(float(sims[i]), 4),
991
+ },
992
+ top5_idx,
993
+ ))
994
+ return {
995
+ "theme_id": t.get("theme_id", 0),
996
+ "theme_label": t.get("theme_label", ""),
997
+ "total_papers": t.get("total_papers", 0),
998
+ "top_5_sentences": top5,
999
+ }
1000
+
1001
+ profiles = list(map(_profile, themes))
1002
+ out_path = Path(themes_json_path).parent / "profiles.json"
1003
+ out_path.write_text(json.dumps(profiles, indent=2))
1004
+
1005
+ return json.dumps({
1006
+ "profiles_count": len(profiles),
1007
+ "profiles_json": str(out_path),
1008
+ "profiles": profiles,
1009
+ "needs_review": True,
1010
+ }, indent=2)
1011
+
1012
+
1013
+ @tool
1014
+ def compare_with_taxonomy(themes_json_path: str) -> str:
1015
+ """Map each theme to PAJAIS 25 IS research categories via Mistral.
1016
+
1017
+ Phase 5.5 β€” Taxonomy Alignment (extension). LLM-DEPENDENT.
1018
+
1019
+ Themes with alignment_score < 0.50 are flagged as potentially NOVEL.
1020
+
1021
+ Args:
1022
+ themes_json_path: Path to themes.json.
1023
+
1024
+ Returns:
1025
+ JSON: themes_aligned count + taxonomy_file path. needs_review=True.
1026
+ """
1027
+ themes = json.loads(Path(themes_json_path).read_text())
1028
+
1029
+ safe_themes = list(map(
1030
+ lambda t: {k: v for k, v in t.items() if k not in ("centroid", "indices")},
1031
+ themes,
1032
+ ))
1033
+
1034
+ template = (
1035
+ "You are an IS research taxonomy expert.\n\n"
1036
+ "PAJAIS 25 Categories:\n{pajais}\n\n"
1037
+ "Research themes:\n{themes}\n\n"
1038
+ "For EACH theme return a JSON array where every element has:\n"
1039
+ " theme_label : string\n"
1040
+ " pajais_categories : list of 1-3 matching PAJAIS category names\n"
1041
+ " alignment_score : float 0.0-1.0\n"
1042
+ " notes : one sentence justification\n\n"
1043
+ "Return ONLY the JSON array β€” no markdown, no preamble."
1044
+ )
1045
+
1046
+ result = _mistral_chain(template).invoke({
1047
+ "pajais": "\n".join(map(lambda c: f"- {c}", PAJAIS_25)),
1048
+ "themes": json.dumps(safe_themes, indent=2),
1049
+ })
1050
+ out_path = Path(themes_json_path).parent / "taxonomy_alignment.json"
1051
+ out_path.write_text(json.dumps(result, indent=2))
1052
+
1053
+ return json.dumps({
1054
+ "themes_aligned": len(result),
1055
+ "taxonomy_file": str(out_path),
1056
+ "needs_review": True,
1057
+ }, indent=2)
1058
+
1059
+
1060
+ @tool
1061
+ def generate_comparison_csv(
1062
+ abstract_themes_path: str,
1063
+ title_themes_path: str,
1064
+ taxonomy_abstract_path: str,
1065
+ taxonomy_title_path: str,
1066
+ ) -> str:
1067
+ """Build side-by-side abstract vs title comparison CSV.
1068
+
1069
+ Phase 6 β€” Report. DETERMINISTIC.
1070
+
1071
+ Joins on PAJAIS_Category. Delta_Score = Abstract - Title.
1072
+
1073
+ Args:
1074
+ abstract_themes_path: themes.json β€” abstract run.
1075
+ title_themes_path: themes.json β€” title run.
1076
+ taxonomy_abstract_path: taxonomy_alignment.json β€” abstract run.
1077
+ taxonomy_title_path: taxonomy_alignment.json β€” title run.
1078
+
1079
+ Returns:
1080
+ JSON: comparison_csv path, total_rows, columns. needs_review=True.
1081
+ """
1082
+ def _explode_taxonomy(path: str) -> pd.DataFrame:
1083
+ """Flatten taxonomy alignment into one row per PAJAIS category."""
1084
+ data = json.loads(Path(path).read_text())
1085
+ rows = sum(
1086
+ list(map(
1087
+ lambda item: list(map(
1088
+ lambda cat: {
1089
+ "pajais_category": cat,
1090
+ "theme_label": item.get("theme_label", ""),
1091
+ "alignment_score": item.get("alignment_score", 0.0),
1092
+ },
1093
+ item.get("pajais_categories", []),
1094
+ )),
1095
+ data,
1096
+ )),
1097
+ [],
1098
+ )
1099
+ return pd.DataFrame(rows)
1100
+
1101
+ df_abs = _explode_taxonomy(taxonomy_abstract_path)
1102
+ df_title = _explode_taxonomy(taxonomy_title_path)
1103
+
1104
+ df_abs.columns = ["PAJAIS_Category", "Abstract_Theme", "Abstract_Score"]
1105
+ df_title.columns = ["PAJAIS_Category", "Title_Theme", "Title_Score"]
1106
+
1107
+ merged = (
1108
+ pd.merge(df_abs, df_title, on="PAJAIS_Category", how="outer")
1109
+ .fillna({"Abstract_Score": 0.0, "Title_Score": 0.0,
1110
+ "Abstract_Theme": "", "Title_Theme": ""})
1111
+ .assign(Delta_Score=lambda d: (d["Abstract_Score"] - d["Title_Score"]).round(4))
1112
+ .sort_values("PAJAIS_Category")
1113
+ .reset_index(drop=True)
1114
+ )
1115
+
1116
+ out_csv = Path(abstract_themes_path).parent / "abstract_vs_title_comparison.csv"
1117
+ merged.to_csv(out_csv, index=False)
1118
+
1119
+ return json.dumps({
1120
+ "comparison_csv": str(out_csv),
1121
+ "total_rows": len(merged),
1122
+ "columns": list(merged.columns),
1123
+ "needs_review": True,
1124
+ }, indent=2)
1125
+
1126
+
1127
+ @tool
1128
+ def export_narrative(
1129
+ taxonomy_alignment_path: str,
1130
+ comparison_csv_path: str,
1131
+ run_mode: str = "abstract",
1132
+ ) -> str:
1133
+ """Generate 500-word Section 7: Discussion & Implications via Mistral.
1134
+
1135
+ Phase 6 β€” Report. LLM-DEPENDENT (grounded in taxonomy + comparison data).
1136
+
1137
+ Args:
1138
+ taxonomy_alignment_path: Path to taxonomy_alignment.json.
1139
+ comparison_csv_path: Path to comparison CSV.
1140
+ run_mode: 'abstract' or 'title'.
1141
+
1142
+ Returns:
1143
+ JSON: narrative_path, word_count, narrative text. needs_review=True.
1144
+ """
1145
+ alignment = json.loads(Path(taxonomy_alignment_path).read_text())
1146
+
1147
+ top_delta = (
1148
+ pd.read_csv(comparison_csv_path)
1149
+ .assign(_abs=lambda d: d["Delta_Score"].abs())
1150
+ .sort_values("_abs", ascending=False)
1151
+ .drop(columns=["_abs"])
1152
+ .head(5)
1153
+ )
1154
+
1155
+ template = (
1156
+ "You are a senior IS researcher writing a systematic literature review.\n\n"
1157
+ "Write Section 7: Discussion & Implications in exactly {word_count} words.\n\n"
1158
+ "Run mode: {run_mode}\n\n"
1159
+ "Taxonomy alignment (top 10):\n{alignment}\n\n"
1160
+ "Top 5 divergent PAJAIS categories (abstract vs title):\n{divergence}\n\n"
1161
+ "Requirements:\n"
1162
+ "1. Discuss dominant themes and PAJAIS alignment.\n"
1163
+ "2. Interpret divergence between abstract- and title-based models.\n"
1164
+ "3. Highlight implications for IS research practice and future agenda.\n"
1165
+ "4. Use formal academic register β€” no bullet points.\n"
1166
+ "5. Return a JSON object with a single key 'narrative' containing the prose.\n\n"
1167
+ "Return ONLY valid JSON."
1168
+ )
1169
+
1170
+ result = _mistral_chain(template).invoke({
1171
+ "word_count": NARRATIVE_WORDS,
1172
+ "run_mode": run_mode,
1173
+ "alignment": json.dumps(alignment[:10], indent=2),
1174
+ "divergence": top_delta.to_json(orient="records", indent=2),
1175
+ })
1176
+ narrative_text = result.get("narrative", str(result))
1177
+ out_path = Path(taxonomy_alignment_path).parent / "narrative.md"
1178
+ out_path.write_text(
1179
+ f"## Section 7: Discussion & Implications\n\n{narrative_text}\n",
1180
+ encoding="utf-8",
1181
+ )
1182
+
1183
+ return json.dumps({
1184
+ "narrative_path": str(out_path),
1185
+ "word_count": len(narrative_text.split()),
1186
+ "narrative": narrative_text,
1187
+ "needs_review": True,
1188
+ }, indent=2)
1189
+
1190
+
1191
+ ALL_TOOLS = [
1192
+ run_phase_1_and_2,
1193
+ load_scopus_csv,
1194
+ run_bertopic_discovery,
1195
+ label_topics_with_llm,
1196
+ reassign_sentences,
1197
+ consolidate_into_themes,
1198
+ compute_saturation,
1199
+ generate_theme_profiles,
1200
+ compare_with_taxonomy,
1201
+ generate_comparison_csv,
1202
+ export_narrative,
1203
+ ]