The short version
- Every candidate page is attributed to a publisher, classified into one of ten types from independent blog through large media.
- The classification is derived from the article itself first. Only an ambiguous result earns extra page fetches, and never more than two.
- Results are cached per workspace and domain for 30 days, so ten articles from one publisher cost one analysis rather than ten.
- The publisher record includes a reachability judgement that feeds the Priority Score — with unknown treated as neutral, not negative.
- Small and independent publishers are kept deliberately. Popularity is not used as a proxy for quality.
Two pages score identically. One is on a solo newsletter where the author writes every issue and answers every reply. The other is on a trade publication with an editorial calendar, a submissions form, and a media kit priced per placement.
The opportunity is the same. The work is not remotely the same. One is an email that gets a reply this week; the other is a process that takes a quarter, if it is open at all. Publisher intelligence is the layer that makes that difference visible before you write anything.
Ten publisher types
Every enriched page is attributed to a publisher, and each publisher is classified:
| Type | What it usually means for outreach |
|---|---|
| Independent blog | One person, direct email, fast decisions. Often the highest reply rate on the whole list. |
| Small company blog | A marketer or founder maintains it. Receptive when the addition genuinely helps their reader. |
| SaaS company blog | A content team with a strategy. Editorial standards are real, and so is the possibility of a competitive conflict. |
| Agency | Publishes on behalf of clients or for its own visibility. Motivations vary; read the page carefully. |
| Newsletter | Curation is the product. Structurally receptive to a good recommendation, and usually reachable directly. |
| Trade publication | Editorial process, contributor guidelines, longer timelines. Worth pursuing, rarely quick. |
| Large media | Formal process, high bar, often a submissions or PR route rather than a personal one. |
| Community | User-generated or forum-shaped. Rules about self-promotion matter more than the pitch. |
| Directory | Listing-driven. Sometimes a legitimate submission, sometimes a paid placement in disguise. |
| Unknown | Not enough signal to classify. Treated as neutral rather than assumed bad. |
Alongside the type, the record captures the publisher's name, its topics, whether it is a company blog, an agency, an independent blog, or a large publication, and a judgement about whether it is plausibly reachable at all — each with a written reason.
Derived from the article, not from a crawl
The obvious way to classify a publisher is to crawl the site. It is also the expensive way, and for a pipeline processing dozens of domains per run it is the difference between a product and a hobby.
So the analysis starts with what has already been paid for: the article itself, which was fetched during enrichment anyway. Most of the time that is enough — an article carries its publisher's voice, its structure, its bylines, and its purpose.
- 01
Analyze the article that is already in hand
No additional fetch. The page's own content, title, and description usually identify the publisher and its type.
- 02
Only if the result is ambiguous, fetch more
When the type comes back unknown or the publisher has no identifiable name, the homepage and the about page are fetched — and that is the whole escalation. Two pages, no site map, no crawl.
- 03
Cache the answer for 30 days
The result is stored per workspace and domain. The next article from that publisher reuses it instead of paying for the analysis again.
Ten articles from one publisher should cost one publisher analysis. Anything else is paying repeatedly for a fact that does not change.
The 30-day freshness window is a deliberate compromise. Publishers do change — a blog gets acquired, a newsletter hires an editor — but they change on a scale of quarters, not days. A month-long cache captures almost all of the savings and almost none of the staleness.
How it feeds the ranking
The publisher's reachability judgement is a direct component of the Reachability Score, which combines with editorial quality to produce the Priority Score. The handling of the unknown case is the part worth spelling out: a publisher known to be reachable scores full marks, one known not to be scores zero, and a publisher with no signal either way scores in the middle.
Treating unknown as failure would have a specific and bad consequence: it would systematically bury every publisher too small to have been catalogued anywhere. Those are frequently the best opportunities on the list. Sitting at neutral keeps them competitive on their editorial merits, which is the only fair way to rank something you genuinely do not know.
Contact discovery order
Runs only after an opportunity qualifies · deduplicated per publisher domain
01The article, already fetched
No extra costAuthor byline and any address published on the page itself.
02Public publisher pages
Fetch onlyHome, /contact, /contact-us, /support, /help, /about, /press, plus one explicitly linked editorial page. Eight pages maximum.
03Provider lookup: the identified author
CreditsA targeted search for the person who wrote the page.
04Provider lookup: domain search
CreditsFallback only, when no author is identified or found.
Verification before the first send. Provider-supplied addresses are verified before outreach goes out. Addresses published on a publisher’s own contact page are not pushed through a verifier by default — several credits per publisher for very little information.
Why small publishers stay on the list
It would be trivial to filter out everything below a traffic threshold, and it would make the list look more impressive. It is not done, for three reasons.
- 1.Relevance concentrates as audience narrows. A newsletter for exactly your buyer is worth more than a general publication that occasionally covers your category.
- 2.Reachability runs the other way to size. The independent blogger reads their own email. The large publication has a form.
- 3.Popularity is not quality. Plenty of high-traffic pages are content farms, and plenty of small ones are the best writing in a niche.
Dice therefore includes small company blogs and independent publishers whenever they are legitimate, relevant, and realistically reachable — and gives you the classification so you can decide for yourself. The one thing it will not do is quietly filter your market down to whatever happens to be popular.