Skip to main content

Our Methodology

How RepVault analyses reputational and risk data

Written by Support Team

RepVault draws on millions of daily documents: news articles, social posts and

other online media, to help leaders understand the external landscape shaping

their reputation and the macro risks influencing their business. This document

explains the motivation, our data model, and the methodology behind it.

Motivation

Executives don't need another feed of headlines. They need to see the shape of

their reputational landscape - what is rising, what is falling, and which of the

forces they care about are driving the change, before deciding where to look

more closely.

Traditional media monitoring starts at the document: it surfaces a stream of

matching articles and leaves you to piece the picture together. RepVault

inverts this. We score every document as it arrives, then present an aggregate

trend view first, so the question a leader answers is "what's changing, and

why?" rather than "which of these thousands of articles matters?". The full

breadth of underlying coverage is always there to drill into but the entry

point is the trend, not the noise.

Crucially, that aggregate view is organised around structures you define. The

topics, drivers, taxonomies and portfolios you build are what turn a generic

media stream into a measure of *your* corporate reputation and *your* risk

exposure, expressed in the terms that matter to your strategy.

Transparency is a core principle. No score in RepVault is a black box.

Every aggregate figure: an Impact total, a sentiment split, a trend line, a

portfolio comparison can be drilled into to reveal exactly which documents it

is derived from, right down to the original source. You never have to take a

number on trust: you can always see what produced it.

Data model

There are a few core objects in the RepVault data model:

- **Documents**: the low-level view of a single news article, social post or

other item that has been ingested and scored by our pipeline. Each carries

its own Impact, sentiment and location signals.

- **Entities**: the subject of analysis. An entity is a **company**, a

**stakeholder**, or a **topic**. Entities provide the scope against which

documents are matched and scored. We hold thousands of companies and

stakeholders, and any that are missing can be added by our expert team.

- **Taxonomies**: a configurable grouping of **Drivers**, where each Driver is a

cluster of **Topics**, and each Topic is defined by specific keywords. A

climate Driver, for example, might gather topics such as climate change,

climate resilience, just transition and renewable energy.

- **Portfolios**: a customisable collection of companies or stakeholders,

created in Settings, used to benchmark against peers or monitor a defined

group. We recommend a maximum of 20 entities per portfolio.

We ingest documents, match and score them against entities and topics, and

aggregate the results through the taxonomy and portfolio structures you

configure. Those aggregates power the Profile, Comparison and Trends views.

Methodology

The scoring pipeline

Documents are scored by a streaming pipeline: as each article or post is

ingested it is enriched and scored in real time, rather than in batches. For

every document, the pipeline derives four signals:

1. **Company significance** — how strongly and centrally the document concerns

the company or stakeholder in question, based on the number and quality of

matches to that entity.

2. **Topic significance** — how strongly the document concerns the topics you

are tracking, based on matches to the keywords, topics and drivers in your

taxonomy.

3. **Reach / credibility** — how far the document is likely to travel and how

trustworthy the source is, combining the **audience reach** of the

publication with its **source credibility**.

4. **Sentiment** — whether coverage is healthy, unhealthy or neutral for the

entity.

5. **Locations** — the geographic locations extracted from the content,

resolved to the regional levels needed for strategic analysis.

Reach / credibility, company significance and topic significance are then

combined into a single **Impact** score for the document.

Impact

Unlike approaches that rank conversations purely by volume, RepVault uses

**Impact** to identify the coverage that genuinely shapes an organisation's

reputation. Impact combines three factors:

> Impact = (Reach / Credibility) × Company significance × Topic significance

Because the three factors are multiplied rather than added, a document only

scores highly when it performs on *all* of them — it must be prominent (high

reach from a credible source), clearly about the company, *and* clearly about a

topic you care about. A well-read, credible article that only mentions your

company in passing, or that concerns your company but not any of your tracked

topics, is scored down accordingly.

Each document receives an Impact score between **0 and 1**. A document from a

highly credible source, reaching a large audience, that is squarely about the

company and one of its tracked topics, approaches 1. A total Impact score for

any view is the sum of the scores of the documents within it — so a view's

Impact reflects both *how much* prominent coverage there is and *how

significant* each piece is, not merely how many articles were published.

This means a single article in a top-tier, high-reach publication that is

squarely about your company and a topic you track can outweigh a large volume of

low-reach, low-significance mentions — which is exactly how reputational

influence works in practice.

Sentiment

Sentiment tells you how coverage about an entity is being perceived. RepVault

classifies each document as **healthy** (positive), **unhealthy** (negative) or

**neutral**, using a widely adopted natural language processing (NLP) standard

trained to predict results across large datasets. Documents the system cannot

confidently classify — primarily non-English content — are marked **unknown**.

As with any NLP approach, there is a margin of accuracy: linguistic features

such as sarcasm, common in social media, are difficult for automated

classification to capture. Online news often reads as neutral because

professional reporting tends to be measured, whereas social content carries more

emotive language and therefore more polarised sentiment.

Locations

Locations are derived in two stages. First, a **geo-location model** reads each

document and extracts *any* location it references — from a city or landmark to

a region or country — wherever it appears in the text. Second, a **geo-parsing

model** resolves those raw mentions into the standardised regional levels needed

for corporate strategic insight, such as country and US state, disambiguating

place names and rolling them up to a consistent geographic hierarchy.

This two-stage approach means a mention of a plant, office or market at street

or city level is still surfaced at the country or regional level a strategy

team works with, rather than being lost because it wasn't stated as a country

outright.

Two independent location signals are captured for each document so you can

distinguish *where a story is about* from *where it is being told*:

- **Location Mentions** — the resolved country or US state referenced within the

content.

- **Publish Location** — the country or US state where the content was

published.

Together these let a global reputational picture be broken down geographically,

separating, for example, coverage originating in a market from coverage that

merely references it.

Taxonomies and portfolios: configuring for your strategy

The scoring pipeline is common to every client; what makes RepVault distinctive

is the structure you build on top of it. This is the core of the platform.

**Topics, Drivers and Taxonomies** let you define exactly which conversations

count and how they roll up. Topics are defined by keywords; Drivers cluster

related topics; Taxonomies group Drivers into a coherent lens on an issue. You

can build your own Taxonomies in Settings, so the categories reflect the risks

and reputational themes specific to your business rather than a generic

industry template. If a topic you need doesn't exist, our expert team can add

it.

**Portfolios** let you define the set of entities you measure against — a peer

group for benchmarking on the Comparison page, or a watch-list of organisations

you are monitoring.

Together, this configurability is what lets RepVault represent an

organisation's strategic priorities as uniquely as possible. The same

underlying data can be shaped into a bespoke measure of corporate reputation

for one client and a macro-risk radar for another, because the drivers,

taxonomies and portfolios — not the raw feed — define what is being measured.

Filtering and drilldown

Once documents are scored and organised, a rich set of post-processing filters

lets you refine the narrative without re-running any analysis:

- **Languages** — restrict the media by language (where more than English is

enabled on your account).

- **Location Mentions** — filter by the country or US state referenced in the

content.

- **Publish Location** — filter by the country or US state of publication.

- **Sentiment** — isolate healthy, unhealthy, neutral or unknown coverage.

- **Media** — choose online media, social media, or (where enabled) your own

connected Sprinklr content.

- **Keywords** — include or exclude specific terms, matching either across the

whole article or on individual references, with exclusions taking priority.

Because filtering happens after scoring, you can move fluidly from the aggregate

picture down to an individual source and back again.

From aggregate trend to source

The Trends, Comparison and Profile views are built to be read top-down. The

Trends view shows how Impact and sentiment for your entities and taxonomies are

moving over time; the Comparison view benchmarks entities and portfolios against

one another; and from any point you can drill into the underlying scored

documents that produced the movement. This is the practical expression of the

motivation above: the same breadth of data as conventional media monitoring, but

surfaced as change and context first, with full provenance one click away.

Reprocessing: keeping pace with a changing world

Entities and language are not static. New companies, stakeholders and topics

emerge; existing ones change name, structure or focus; and the vocabulary used

to discuss an issue shifts, often rapidly, once it comes under the media

spotlight. A definition that captured a topic accurately last quarter can drift

out of date as new terms, actors and framings enter the conversation.

RepVault is built to evolve with this. When a new entity is added, or an

existing entity or topic definition is created or changed, we can **reprocess**

previously ingested documents against the updated definitions — re-scoring the

back catalogue rather than only applying changes to new coverage. This means a

newly added company, or an updated set of keywords for a fast-moving topic, is

reflected consistently across history as well as going forward.

This reprocessing capability is what lets us keep language matching current for

topics that evolve under scrutiny, and ensures that when your taxonomy changes,

the resulting measures remain comparable over time rather than breaking at the

point of the change.

Human expertise and continual improvement

Automated scoring is complemented by human expertise. Missing companies,

stakeholders or topics can be requested and added by our team, and taxonomies

and portfolios remain under your editorial control at all times.

We refine our methodology regularly to incorporate new research and technical

capability, and we welcome feedback from industry and academic colleagues. To

learn more or work with us on future development, please contact

Did this answer your question?