What do local party organizations do and say? CLARA measures the
observable web presence of 2,635 local chapters of the seven main
German parties (CDU, CSU, SPD, Grüne, Linke, AfD, FDP), annually from
2015 to 2025: 28,985 chapter-years in total. The unit of observation is
the county-level party association, the Kreisverband, of which each
party has one per county. Of the 2,635 chapters, 2,591 are
Kreisverbände, as are 2,316 of the 2,319 chapters in the recommended
analysis scope. The rest are city- or town-level chapters, which the
org_level column identifies. The public release contains
chapter-year aggregates and coverage metadata only: no personal data,
no reproduced website text.
How the data were collected
The sampling frame is fixed before any website discovery: 2,394
party × county cells, one cell for each party that organizes in a
given county. The official municipality register, restricted to
municipalities with at least 5,000 residents, covers 399 Kreise
(counties). Each Kreis is crossed with the five parties that organize
everywhere, plus the CDU outside Bavaria and the CSU within it. Each
chapter is georeferenced to its seat municipality, the town the
association is based in, through the 8-digit municipality key
(ags), so the panel merges
with election and demographic data. While a Kreisverband organizes the
whole county, covariates attached to the seat municipality describe
only that town; merges that need county-wide covariates should use the
5-digit county key (kreis_id). A Kreisverband typically
contains subordinate town-level chapters (Ortsverbände, Stadtverbände),
which are not separate observations. Where a subordinate chapter's
pages sit under the Kreisverband's domain, collection stays inside each
chapter's own URL path, so those pages are not attributed to the
Kreisverband. Chapter discovery located a website for 98.8% of cells
(2,365 of 2,394) through official party directories, a national roster PDF,
systematic URL probing, search-engine queries whose results a language
model screened, and, for sites that had already gone offline, archived
copies of the party directories; the url_source column
records which of these found each chapter. No website could be found
for 27 cells (14 Linke, 6 FDP, 3 AfD, 2 SPD, 1 CDU, 1 Grüne), and 2
further chapters are known but have no usable URL.
Each chapter's website is then reconstructed year by year from
Internet Archive Wayback Machine snapshots, roughly 1.03 million
archived pages in total. A locally hosted open-weight language model
extracts structured records from these pages
(Qwen3-30B-A3B-Instruct-2507, FP8, served with vLLM), and
schema-constrained decoding restricts its output to the fields of the
target record. The page-function labels that route pages to extractors
combine URL rules with the same local model, and every model-derived
value in the release comes from the local
model. One measure is built without a model: where a chapter has no
archived about or contact page, its social-media accounts come from
scanning the stored page text for links to the six social-media
platforms the dataset covers, and the social_source column
marks those records. News pages are
deduplicated to one record per article (237,399 articles dated within
the panel window) and
classified by policy topic, article function, and geographic focus. The
data report describes each step.
What the data measure
- Organizational presence: whether a leadership board, officials, issue content, and a local office address are listed.
- Board composition: size, gender shares inferred from first names, whether a chair is listed and that chair's gender, and proxies for professionalization (the shares of officials with an academic title, a listed email address, or a photo).
- Issue emphasis: 25 issue-topic families plus an unclassified remainder, each family mapped to a major topic code of the Comparative Agendas Project so the measures line up with other agenda-coded data; how widely a chapter spreads its content across families; and how much of that content falls in each family.
- Digital presence: chapter-specific Facebook, Instagram, X/Twitter, YouTube, Telegram, and TikTok accounts (national party accounts excluded).
- Mobilization infrastructure: flags for whether the site offers a donation, membership, newsletter, program, events, or volunteer pathway. These cover the information, participation, and mobilization functions in Gibson and Ward's scheme for coding party websites.
- News content: deduplicated article counts plus shares by policy topic, article function (press release, event, position statement, campaign, council work), and geographic focus (local / regional / national).
Caveats
Coverage. Of the 25,509 chapter-years in the recommended
analysis scope, 17,989 (70.5%) are researcher-usable, meaning at least
one measure was observed for them. Each panel row records what its
coverage rests on (coverage_basis, plus the finer
release_coverage_bucket and
release_coverage_tier labels), so a user can rebuild the
same denominator or impose a stricter one. Roughly 25% of chapter-years
have no same-year archived snapshot in the Internet Archive, so no
further collection can recover them. Where the archive does hold a page
of the type a measure needs, extraction coverage is 99.9% for
officials, 100% for contact and social measures, 99.7% for issue
positions, and 92.7% for news. Coverage rises over time and varies by
party, so analyses should condition on the coverage columns. The
deposit also includes an audit that assigns every missing value a cause
and estimates how much extractable content the archive holds but the
release does not capture.
Collection and extraction. Chapter discovery ran against the
live web in early 2026, so chapters whose sites went offline earlier
without replacement are missing entirely. The early panel years
therefore carry survivorship selection on top of archive availability: a
chapter had to survive to 2026 to enter the roster at all, and its site
had to have been archived in a given year for the chapter to be observed
that year. A chapter's news count depends on how often
the Internet Archive captured its site, not only on how much the chapter
posted. The board measures mix two groups of people, internal party
officers and elected council representatives; the
officials_source column reports which page type a
chapter-year's board came from, so analyses can condition on it. Gender
is inferred from first names, a procedure that is correct about 94% of
the time. The issue fields record which topics a chapter writes about,
not what stance it takes.
Before analysis. Read ags and
kreis_id as text rather than as numbers: both keys carry
leading zeros, which a numeric import silently strips.
kreis_id is the 5-digit county key and the recommended key
for merging county-level data; ags is the 8-digit key of
the chapter's seat municipality. To reproduce the analysis sample, keep
the rows with analysis_scope_included = 1, then pick a
coverage rule. The conservative choice,
coverage_basis = 'full', keeps only the chapter-years that
meet the structured-coverage threshold; any non-blank
coverage_basis also keeps rows whose values were observed
on thinner or fallback pages.
Validation
Three kinds of evidence bear on whether the extracted values are right. First, a human coder reviewed extracted records against excerpts of the pages they came from, across five stratified samples. The first 288 records reviewed all matched. On a 48-record sample drawn from the released records themselves, 45 were confirmed, 1 was incorrect, and 2 were undecidable from the excerpt. Second, an independent locally hosted model from a different family than the extractor (Mistral-Small-24B) judged a fresh 315-record sample of the released records against their source pages. Of 769 judged non-empty fields, 76.9% were rated supported and 2.9% contradicted; the rest did not appear in the truncated, navigation-stripped excerpt the judge was shown and are therefore neither confirmed nor refuted. The third kind of evidence is face validity: the three figures below show that the news labels move with real-world events the classifier is never told about.
The share of articles labeled election campaigning rises and falls with the electoral calendar, even though the classifier never sees the year.
Policy topics behave the same way. Attention to migration, climate, health, and energy each peaks inside or at the edge of the period in which that issue dominated German national debate.
Migration coverage also separates the parties in the expected direction: the AfD devotes a larger share of its classified news topics to migration than any other party, in every year of the panel.
Data and documentation
| File | Grain | Rows |
|---|---|---|
chapter_roster.csv | chapter | 2,635 |
chapter_year_panel.csv | chapter-year | 28,985 |
chapter_year_panel_enriched.csv | chapter-year | 28,985 |
chapter_year_topic.csv | chapter-year-family | 39,671 |
chapter_year_news.csv | chapter-year | 14,733 |
chapter_year_news_topic.csv | chapter-year-family | 90,668 |
chapter_year_provenance.csv | chapter-year | 28,985 |
topic_family_cap_crosswalk.csv | family | 26 |
kreis_wahlkreis_crosswalk.csv | county-election-district | 1,640 |
The deposit bundle also contains the codebook, the data report, coverage and missingness diagnostics, validation summaries, and minimal R and Python load scripts. The core panel is mirrored here as chapter_year_panel.csv (13 MB), but the Harvard Dataverse deposit is the citable version. Person-level records and reproduced page text are not part of the release: the records are personal data under the GDPR, and the text may be protected by copyright. A controlled-access version of that material is in preparation and is not ready yet. Data and documentation are CC-BY-4.0, and the pipeline code is available from the author on request.
Citation
Hilbig, Hanno (2026). CLARA: Communication of Local German Party Associations, Recorded Annually (2015-2025). Version 1.0.0. Harvard Dataverse. https://doi.org/10.7910/DVN/2KJSCO