RepVault draws on millions of daily documents: news articles, social posts and
other online media, to help leaders understand the external landscape shaping
their reputation and the macro risks influencing their business. This document
explains the motivation, our data model, and the methodology behind it.
Motivation
Executives don't need another feed of headlines. They need to see the shape of
their reputational landscape - what is rising, what is falling, and which of the
forces they care about are driving the change, before deciding where to look
more closely.
Traditional media monitoring starts at the document: it surfaces a stream of
matching articles and leaves you to piece the picture together. RepVault
inverts this. We score every document as it arrives, then present an aggregate
trend view first, so the question a leader answers is "what's changing, and
why?" rather than "which of these thousands of articles matters?". The full
breadth of underlying coverage is always there to drill into but the entry
point is the trend, not the noise.
Crucially, that aggregate view is organised around structures you define. The
topics, drivers, taxonomies and portfolios you build are what turn a generic
media stream into a measure of *your* corporate reputation and *your* risk
exposure, expressed in the terms that matter to your strategy.
Transparency is a core principle. No score in RepVault is a black box.
Every aggregate figure: an Impact total, a sentiment split, a trend line, a
portfolio comparison can be drilled into to reveal exactly which documents it
is derived from, right down to the original source. You never have to take a
number on trust: you can always see what produced it.
Data model
There are a few core objects in the RepVault data model:
- **Documents**: the low-level view of a single news article, social post or
other item that has been ingested and scored by our pipeline. Each carries
its own Impact, sentiment and location signals.
- **Entities**: the subject of analysis. An entity is a **company**, a
**stakeholder**, or a **topic**. Entities provide the scope against which
documents are matched and scored. We hold thousands of companies and
stakeholders, and any that are missing can be added by our expert team.
- **Taxonomies**: a configurable grouping of **Drivers**, where each Driver is a
cluster of **Topics**, and each Topic is defined by specific keywords. A
climate Driver, for example, might gather topics such as climate change,
climate resilience, just transition and renewable energy.
- **Portfolios**: a customisable collection of companies or stakeholders,
created in Settings, used to benchmark against peers or monitor a defined
group. We recommend a maximum of 20 entities per portfolio.
We ingest documents, match and score them against entities and topics, and
aggregate the results through the taxonomy and portfolio structures you
configure. Those aggregates power the Profile, Comparison and Trends views.
Methodology
The scoring pipeline
Documents are scored by a streaming pipeline: as each article or post is
ingested it is enriched and scored in real time, rather than in batches. For
every document, the pipeline derives four signals:
1. **Company significance** — how strongly and centrally the document concerns
the company or stakeholder in question, based on the number and quality of
matches to that entity.
2. **Topic significance** — how strongly the document concerns the topics you
are tracking, based on matches to the keywords, topics and drivers in your
taxonomy.
3. **Reach / credibility** — how far the document is likely to travel and how
trustworthy the source is, combining the **audience reach** of the
publication with its **source credibility**.
4. **Sentiment** — whether coverage is healthy, unhealthy or neutral for the
entity.
5. **Locations** — the geographic locations extracted from the content,
resolved to the regional levels needed for strategic analysis.
Reach / credibility, company significance and topic significance are then
combined into a single **Impact** score for the document.
Impact
Unlike approaches that rank conversations purely by volume, RepVault uses
**Impact** to identify the coverage that genuinely shapes an organisation's
reputation. Impact combines three factors:
> Impact = (Reach / Credibility) × Company significance × Topic significance
Because the three factors are multiplied rather than added, a document only
scores highly when it performs on *all* of them — it must be prominent (high
reach from a credible source), clearly about the company, *and* clearly about a
topic you care about. A well-read, credible article that only mentions your
company in passing, or that concerns your company but not any of your tracked
topics, is scored down accordingly.
Each document receives an Impact score between **0 and 1**. A document from a
highly credible source, reaching a large audience, that is squarely about the
company and one of its tracked topics, approaches 1. A total Impact score for
any view is the sum of the scores of the documents within it — so a view's
Impact reflects both *how much* prominent coverage there is and *how
significant* each piece is, not merely how many articles were published.
This means a single article in a top-tier, high-reach publication that is
squarely about your company and a topic you track can outweigh a large volume of
low-reach, low-significance mentions — which is exactly how reputational
influence works in practice.
Sentiment
Sentiment tells you how coverage about an entity is being perceived. RepVault
classifies each document as **healthy** (positive), **unhealthy** (negative) or
**neutral**, using a widely adopted natural language processing (NLP) standard
trained to predict results across large datasets. Documents the system cannot
confidently classify — primarily non-English content — are marked **unknown**.
As with any NLP approach, there is a margin of accuracy: linguistic features
such as sarcasm, common in social media, are difficult for automated
classification to capture. Online news often reads as neutral because
professional reporting tends to be measured, whereas social content carries more
emotive language and therefore more polarised sentiment.
Locations
Locations are derived in two stages. First, a **geo-location model** reads each
document and extracts *any* location it references — from a city or landmark to
a region or country — wherever it appears in the text. Second, a **geo-parsing
model** resolves those raw mentions into the standardised regional levels needed
for corporate strategic insight, such as country and US state, disambiguating
place names and rolling them up to a consistent geographic hierarchy.
This two-stage approach means a mention of a plant, office or market at street
or city level is still surfaced at the country or regional level a strategy
team works with, rather than being lost because it wasn't stated as a country
outright.
Two independent location signals are captured for each document so you can
distinguish *where a story is about* from *where it is being told*:
- **Location Mentions** — the resolved country or US state referenced within the
content.
- **Publish Location** — the country or US state where the content was
published.
Together these let a global reputational picture be broken down geographically,
separating, for example, coverage originating in a market from coverage that
merely references it.
Taxonomies and portfolios: configuring for your strategy
The scoring pipeline is common to every client; what makes RepVault distinctive
is the structure you build on top of it. This is the core of the platform.
**Topics, Drivers and Taxonomies** let you define exactly which conversations
count and how they roll up. Topics are defined by keywords; Drivers cluster
related topics; Taxonomies group Drivers into a coherent lens on an issue. You
can build your own Taxonomies in Settings, so the categories reflect the risks
and reputational themes specific to your business rather than a generic
industry template. If a topic you need doesn't exist, our expert team can add
it.
**Portfolios** let you define the set of entities you measure against — a peer
group for benchmarking on the Comparison page, or a watch-list of organisations
you are monitoring.
Together, this configurability is what lets RepVault represent an
organisation's strategic priorities as uniquely as possible. The same
underlying data can be shaped into a bespoke measure of corporate reputation
for one client and a macro-risk radar for another, because the drivers,
taxonomies and portfolios — not the raw feed — define what is being measured.
Filtering and drilldown
Once documents are scored and organised, a rich set of post-processing filters
lets you refine the narrative without re-running any analysis:
- **Languages** — restrict the media by language (where more than English is
enabled on your account).
- **Location Mentions** — filter by the country or US state referenced in the
content.
- **Publish Location** — filter by the country or US state of publication.
- **Sentiment** — isolate healthy, unhealthy, neutral or unknown coverage.
- **Media** — choose online media, social media, or (where enabled) your own
connected Sprinklr content.
- **Keywords** — include or exclude specific terms, matching either across the
whole article or on individual references, with exclusions taking priority.
Because filtering happens after scoring, you can move fluidly from the aggregate
picture down to an individual source and back again.
From aggregate trend to source
The Trends, Comparison and Profile views are built to be read top-down. The
Trends view shows how Impact and sentiment for your entities and taxonomies are
moving over time; the Comparison view benchmarks entities and portfolios against
one another; and from any point you can drill into the underlying scored
documents that produced the movement. This is the practical expression of the
motivation above: the same breadth of data as conventional media monitoring, but
surfaced as change and context first, with full provenance one click away.
Reprocessing: keeping pace with a changing world
Entities and language are not static. New companies, stakeholders and topics
emerge; existing ones change name, structure or focus; and the vocabulary used
to discuss an issue shifts, often rapidly, once it comes under the media
spotlight. A definition that captured a topic accurately last quarter can drift
out of date as new terms, actors and framings enter the conversation.
RepVault is built to evolve with this. When a new entity is added, or an
existing entity or topic definition is created or changed, we can **reprocess**
previously ingested documents against the updated definitions — re-scoring the
back catalogue rather than only applying changes to new coverage. This means a
newly added company, or an updated set of keywords for a fast-moving topic, is
reflected consistently across history as well as going forward.
This reprocessing capability is what lets us keep language matching current for
topics that evolve under scrutiny, and ensures that when your taxonomy changes,
the resulting measures remain comparable over time rather than breaking at the
point of the change.
Human expertise and continual improvement
Automated scoring is complemented by human expertise. Missing companies,
stakeholders or topics can be requested and added by our team, and taxonomies
and portfolios remain under your editorial control at all times.
We refine our methodology regularly to incorporate new research and technical
capability, and we welcome feedback from industry and academic colleagues. To
learn more or work with us on future development, please contact