Monitoring emerging and disruptive technologies (EDTs) is important for anticipating developments that may affect national cyber-defence capabilities, technological sovereignty, critical infrastructure, and strategic dependencies. Relevant information is distributed across scientific publications, patents, standards, technical reports, procurement notices, and industry announcements.

Large language models could support this process by extracting technologies, organizations, capabilities, maturity indicators, and relationships from large document collections. They could also synthesize evidence into concise technology assessments. However, different LLMs may vary considerably in accuracy, consistency, evidence grounding, computational cost, privacy, and suitability for deployment in sensitive environments.

Current LLM benchmarks primarily evaluate general knowledge, reasoning, coding, and language understanding. These benchmarks do not adequately measure the capabilities required for strategic technology monitoring. A model that performs well on general benchmarks may still struggle to distinguish demonstrated technical progress from speculation, assess technology maturity, or provide traceable evidence.

This thesis will develop a task-specific benchmark for evaluating commercial and open-source LLMs in the monitoring of EDTs relevant to national cyber defence.

Objectives

The objective is to design and apply a benchmark that systematically evaluates the capabilities and limitations of different LLMs for evidence-grounded EDT monitoring.

The study will examine not only model accuracy but also consistency, explainability, efficiency, privacy, and suitability for local deployment.

Research questions :

  • How accurately do different LLMs identify and classify emerging technologies?
  • Which models perform best at extracting actors, capabilities, applications, limitations, and maturity indicators?
  • Can the models distinguish demonstrated progress from speculation, marketing claims, and technological hype?
  • How accurately do the models connect generated conclusions to supporting evidence?
  • Does RAG improve all models equally?
  • How consistent are model outputs across prompts, repeated runs, and technology domains?
  • What trade-offs exist between accuracy, latency, cost, privacy, and hardware requirements?
  • Can smaller, locally deployable models provide sufficient performance for sensitive government applications?

Expected contributions :

The thesis should produce:

  • A reusable benchmark for LLM-based EDT monitoring.
  • An annotated dataset covering selected technology domains.
  • A systematic comparison of commercial and open-source models.
  • Evidence about when RAG improves monitoring performance.
  • A model-selection framework considering performance, cost, privacy, and deployment constraints.
  • Recommendations for using LLMs in sensitive technology-monitoring environments.