CLARA

Communication of Local German Party Associations, Recorded Annually
A chapter-year panel dataset, 2015–2025 · Version 1.0.0 · Hanno Hilbig, University of California, Davis

What do local party organizations do and say? CLARA measures the observable web presence of 2,635 local chapters of the seven main German parties (CDU, CSU, SPD, Grüne, Linke, AfD, FDP), annually from 2015 to 2025: 28,985 chapter-years in total. The unit of observation is the county-level party association, the Kreisverband, of which each party has one per county. Of the 2,635 chapters, 2,591 are Kreisverbände, as are 2,316 of the 2,319 chapters in the recommended analysis scope. The rest are city- or town-level chapters, which the org_level column identifies. The public release contains chapter-year aggregates and coverage metadata only: no personal data, no reproduced website text.

How the data were collected

The sampling frame is fixed before any website discovery: 2,394 party × county cells, one cell for each party that organizes in a given county. The official municipality register, restricted to municipalities with at least 5,000 residents, covers 399 Kreise (counties). Each Kreis is crossed with the five parties that organize everywhere, plus the CDU outside Bavaria and the CSU within it. Each chapter is georeferenced to its seat municipality, the town the association is based in, through the 8-digit municipality key (ags), so the panel merges with election and demographic data. While a Kreisverband organizes the whole county, covariates attached to the seat municipality describe only that town; merges that need county-wide covariates should use the 5-digit county key (kreis_id). A Kreisverband typically contains subordinate town-level chapters (Ortsverbände, Stadtverbände), which are not separate observations. Where a subordinate chapter's pages sit under the Kreisverband's domain, collection stays inside each chapter's own URL path, so those pages are not attributed to the Kreisverband. Chapter discovery located a website for 98.8% of cells (2,365 of 2,394) through official party directories, a national roster PDF, systematic URL probing, search-engine queries whose results a language model screened, and, for sites that had already gone offline, archived copies of the party directories; the url_source column records which of these found each chapter. No website could be found for 27 cells (14 Linke, 6 FDP, 3 AfD, 2 SPD, 1 CDU, 1 Grüne), and 2 further chapters are known but have no usable URL.

Each chapter's website is then reconstructed year by year from Internet Archive Wayback Machine snapshots, roughly 1.03 million archived pages in total. A locally hosted open-weight language model extracts structured records from these pages (Qwen3-30B-A3B-Instruct-2507, FP8, served with vLLM), and schema-constrained decoding restricts its output to the fields of the target record. The page-function labels that route pages to extractors combine URL rules with the same local model, and every model-derived value in the release comes from the local model. One measure is built without a model: where a chapter has no archived about or contact page, its social-media accounts come from scanning the stored page text for links to the six social-media platforms the dataset covers, and the social_source column marks those records. News pages are deduplicated to one record per article (237,399 articles dated within the panel window) and classified by policy topic, article function, and geographic focus. The data report describes each step.

What the data measure

Caveats

Coverage. Of the 25,509 chapter-years in the recommended analysis scope, 17,989 (70.5%) are researcher-usable, meaning at least one measure was observed for them. Each panel row records what its coverage rests on (coverage_basis, plus the finer release_coverage_bucket and release_coverage_tier labels), so a user can rebuild the same denominator or impose a stricter one. Roughly 25% of chapter-years have no same-year archived snapshot in the Internet Archive, so no further collection can recover them. Where the archive does hold a page of the type a measure needs, extraction coverage is 99.9% for officials, 100% for contact and social measures, 99.7% for issue positions, and 92.7% for news. Coverage rises over time and varies by party, so analyses should condition on the coverage columns. The deposit also includes an audit that assigns every missing value a cause and estimates how much extractable content the archive holds but the release does not capture.

Small multiples: usable structured coverage by year for each of the seven parties, rising over time with clear party differences.
Usable structured coverage by party and year.

Collection and extraction. Chapter discovery ran against the live web in early 2026, so chapters whose sites went offline earlier without replacement are missing entirely. The early panel years therefore carry survivorship selection on top of archive availability: a chapter had to survive to 2026 to enter the roster at all, and its site had to have been archived in a given year for the chapter to be observed that year. A chapter's news count depends on how often the Internet Archive captured its site, not only on how much the chapter posted. The board measures mix two groups of people, internal party officers and elected council representatives; the officials_source column reports which page type a chapter-year's board came from, so analyses can condition on it. Gender is inferred from first names, a procedure that is correct about 94% of the time. The issue fields record which topics a chapter writes about, not what stance it takes.

Before analysis. Read ags and kreis_id as text rather than as numbers: both keys carry leading zeros, which a numeric import silently strips. kreis_id is the 5-digit county key and the recommended key for merging county-level data; ags is the 8-digit key of the chapter's seat municipality. To reproduce the analysis sample, keep the rows with analysis_scope_included = 1, then pick a coverage rule. The conservative choice, coverage_basis = 'full', keeps only the chapter-years that meet the structured-coverage threshold; any non-blank coverage_basis also keeps rows whose values were observed on thinner or fallback pages.

Validation

Three kinds of evidence bear on whether the extracted values are right. First, a human coder reviewed extracted records against excerpts of the pages they came from, across five stratified samples. The first 288 records reviewed all matched. On a 48-record sample drawn from the released records themselves, 45 were confirmed, 1 was incorrect, and 2 were undecidable from the excerpt. Second, an independent locally hosted model from a different family than the extractor (Mistral-Small-24B) judged a fresh 315-record sample of the released records against their source pages. Of 769 judged non-empty fields, 76.9% were rated supported and 2.9% contradicted; the rest did not appear in the truncated, navigation-stripped excerpt the judge was shown and are therefore neither confirmed nor refuted. The third kind of evidence is face validity: the three figures below show that the news labels move with real-world events the classifier is never told about.

The share of articles labeled election campaigning rises and falls with the electoral calendar, even though the classifier never sees the year.

Bar chart: share of local party news articles classified as campaign content by year, peaking in German federal and European election years.
Share of news articles classified as election campaigning, by year (BTW = federal election, EP = European election).

Policy topics behave the same way. Attention to migration, climate, health, and energy each peaks inside or at the edge of the period in which that issue dominated German national debate.

Four line charts: the share of classified news topics devoted to integration and migration, environment and climate, health, and energy, by year from 2015 to 2025, each peaking during or at the edge of the shaded period when that issue was nationally salient.
Share of classified news topics by year for four policy families. The shaded bands mark the 2015–16 refugee crisis, the 2019 climate protests, the 2020–21 pandemic, and the 2022 energy crisis; the labeled point in each panel is that family's maximum.

Migration coverage also separates the parties in the expected direction: the AfD devotes a larger share of its classified news topics to migration than any other party, in every year of the panel.

Line chart: share of news topics about integration and migration by year, with the AfD line above the lines for all other parties throughout, peaking in 2015 and rising again around the 2023 migration debate.
Share of classified news topics in the integration and migration family, by year. The dark line is the AfD, the grey lines the other parties; the shaded bands mark the 2015–16 refugee crisis and the 2023 migration debate.

Data and documentation

FileGrainRows
chapter_roster.csvchapter2,635
chapter_year_panel.csvchapter-year28,985
chapter_year_panel_enriched.csvchapter-year28,985
chapter_year_topic.csvchapter-year-family39,671
chapter_year_news.csvchapter-year14,733
chapter_year_news_topic.csvchapter-year-family90,668
chapter_year_provenance.csvchapter-year28,985
topic_family_cap_crosswalk.csvfamily26
kreis_wahlkreis_crosswalk.csvcounty-election-district1,640

The deposit bundle also contains the codebook, the data report, coverage and missingness diagnostics, validation summaries, and minimal R and Python load scripts. The core panel is mirrored here as chapter_year_panel.csv (13 MB), but the Harvard Dataverse deposit is the citable version. Person-level records and reproduced page text are not part of the release: the records are personal data under the GDPR, and the text may be protected by copyright. A controlled-access version of that material is in preparation and is not ready yet. Data and documentation are CC-BY-4.0, and the pipeline code is available from the author on request.

Citation

Hilbig, Hanno (2026). CLARA: Communication of Local German
Party Associations, Recorded Annually (2015-2025). Version 1.0.0.
Harvard Dataverse. https://doi.org/10.7910/DVN/2KJSCO