# Licences for the data files in this directory

Everything in `public/data/` is served at a public URL. Putting a file here is
redistribution, whether or not a page draws it. This file records, for each one,
where it came from, what licence it is under, when it was taken, and the exact
text a consumer must reproduce.

Compiled **2026-09-30**. Every byte count, feature count and SHA-256 below was
measured on that date on this machine. Every licence statement was read on that
date from the source named, not recalled. Upstream repositories were read with
read-only HTTP GETs.

The machine-readable copy of the attribution strings is `src/lib/geoAttribution.ts`.
The full provenance record, including defects and export rules, is
`docs/data-provenance.md`; build notes for the two constituency layers are in
`docs/geography-sources.md`.

---

## At a glance

| File | Bytes | Features | Licence | Attribution |
|---|---|---|---|---|
| `pc-boundaries.geojson` | 1,539,633 | 543 | CC0 1.0 | Courtesy |
| `ac-haryana.geojson` | 170,281 | 90 | CC BY 2.5 India | **Required** |
| `india-districts.geojson` | 3,153,362 | 742 | CC0 1.0 | Requested |
| `district-meta.json` | 58,776 | 742 keys | CC0 1.0 (derived from the above) | Requested |

This file sits in `public/data/` so that it deploys to
`https://interns.city/data/LICENSES.md`, alongside the files it describes — a
consumer who downloads one of them can reach its licence from the same directory.
It is unpublished as of 2026-09-30, because this change is not deployed.

---

## 1. `pc-boundaries.geojson`

Parliamentary constituency boundaries — 543 Lok Sabha seats, national.

| | |
|---|---|
| Source URL | `https://raw.githubusercontent.com/datameet/maps/master/parliamentary-constituencies/india_pc_2019_simplified.geojson` |
| Upstream | DataMeet / Arun Ganesh |
| Licence | **CC0 1.0 Universal Public Domain Dedication** |
| Licence URL | `https://creativecommons.org/publicdomain/zero/1.0/` |
| Fetched | 2026-09-29 |
| Source SHA-256 | `54840686c3c5ceabb223940127f3067d921d052abd312e2eb4c6784c323ffc41` |
| Our file SHA-256 | `080926e8808c42820e50d9b8741e8e1d889190fc7d52d7e3007cdf17eea095a2` |
| Public export | Yes, unconditionally |

**Attribution text.** CC0 waives it. Reproduce it anyway so a reader can trace the
data:

> Parliamentary constituency boundaries: Arun Ganesh / DataMeet
> (india_pc_2019_simplified.geojson, CC0 1.0). Attribution is courtesy, not a
> licence term.

That sentence is the file's own `attribution` member, quoted verbatim.

**Licence trap.** The same DataMeet directory also holds `india_pc_2019.shp`, a
48.8 MB full-resolution shapefile, which is **CC BY-SA 2.5 India — share-alike**.
It is the obvious upgrade when someone wants crisper coastlines, and swapping it
in would make every derived extract share-alike. GitHub's API reports the whole
`datameet/maps` repository as MIT, which is wrong for both files. Three licences,
one repository, per-directory READMEs. Record the licence per file.

---

## 2. `ac-haryana.geojson`

Assembly constituency boundaries, Haryana only — 90 of the 4,182 polygons in the
national source file.

| | |
|---|---|
| Source URL | `https://raw.githubusercontent.com/datameet/maps/master/assembly-constituencies/India_AC.shp` (plus `.dbf`, `.shx`, `.prj`) |
| Upstream | DataMeet, scraped from the Election Commission of India's polling-station website |
| Licence | **Creative Commons Attribution 2.5 India** |
| Licence URL | `http://creativecommons.org/licenses/by/2.5/in/` |
| Fetched | 2026-09-29 |
| Our file SHA-256 | `719ad4d1954f577c73b576c3cd5505ab94025f67e568def2fb469efb8576aa9b` |
| Public export | Yes, provided the attribution below travels with it |

**Attribution text — this is a licence term, not a courtesy.** Reproduce it
verbatim wherever this layer is rendered:

> Assembly constituency boundaries: DataMeet, India Assembly Constituencies
> (India_AC), scraped from the Election Commission of India's polling-station
> website, shared under Creative Commons Attribution 2.5 India.

That sentence is the file's own `attribution` member, quoted verbatim. The file
also sets `"attribution_required": true`.

**Where the obligation is currently discharged.** `/representatives/near` renders
it. Verified 2026-09-30 by a read-only GET of
`https://interns.city/representatives/near?lat=28.4595&lng=77.0266` (HTTP 200):
the response body contains the string *Creative Commons Attribution 2.5 India*
and the sentence above. Any future page that draws this layer must carry it too —
import it from `src/lib/geoAttribution.ts` rather than retyping it.

---

## 3. `india-districts.geojson`

District boundaries — 742 polygons, national. The most widely consumed geography
in the codebase, and until now the only file here with no licence notice at all.

### The licence, established

**CC0 1.0 Universal Public Domain Dedication**, with attribution *requested* by
the republisher rather than required.

Three statements about this file appeared to conflict. They do not, once each is
read at the level it actually applies to. The chain, every link read 2026-09-30:

| Statement | What it actually governs |
|---|---|
| The `india-geodata` README badge says `CC BY 4.0` | The repository. Its own `LICENSE` file settles this: *"The CC BY 4.0 license above applies to the repository structure, documentation, and any original work. Individual datasets carry their own licenses as documented in their respective metadata.json files."* The badge was never a claim about this file. |
| GitHub's API reports `spdx_id: "NOASSERTION"` | Nothing. The repository's `LICENSE` is hand-written prose, so GitHub's licence detector could not match it to an SPDX identifier. That is a detector result, not a licence statement. |
| `data/administrative/districts/metadata.json` says `"CC0-1.0 / CC-BY-4.0"` | Both, for different files — and the directory's own `README.md`, one level more specific, says which is which: **"Release files: CC0 (Public Domain). DataMeet files: CC BY 4.0."** `SOI_Districts.parquet`, the file we took, is a release file. |

### The derivation, proven rather than assumed

`yashveeeeeeer/india-geodata` release `admin/districts` (published
2026-03-08T04:56:31Z) is a mirror of `ramSeraph/indian_admin_boundaries` release
`districts` (published 2023-12-11T10:14:05Z):

- Six of six non-Parquet assets are identical in byte size across the two
  releases (`SOI_Districts.geojsonl.7z` 27,701,204; `SOI_Districts.pmtiles`
  12,395,728; and the LGD and Bhuvan equivalents).
- For `SOI_Districts.geojsonl.7z`, ranged GETs of the first 65,536 bytes
  (SHA-256 `060f3c910797b1dff70cacbdfdb52b5e1e4a17803a3874136a9a2e2355e3f450`) and
  of the last 65,536 bytes (SHA-256
  `1f7729f577e8d842b477ff15d2022434c3deeff8ffcdbace9005ac6cff2095b9`) match
  byte-for-byte between the two releases.
- Only the `.parquet` assets differ in size (28,691,382 against 46,160,710) —
  consistent with a mirror that re-encoded Parquet and copied everything else.

That matters because `ramSeraph`, the origin, states the licence **per file**, in
the release notes for this exact asset:

> Source: Survey Of India — `https://onlinemaps.surveyofindia.gov.in/Digital_Product_Show.aspx`
> License: CC0 1.0 but attribute datameet and the original government source where possible

linking to `https://github.com/ramSeraph/indianopenmaps/blob/main/DATA_LICENSE.md`,
which says, verbatim:

> Most of the data here (anything that links here for licensing) is licensed using
> CC0 1.0. This means that I am relinquishing whatever rights I might have on this
> data. However, I request that you actively acknowledge and give attribution to
> Datameet community for collating the data and also the original government
> sources wherever possible.

### What is still not established, stated plainly

CC0 here is a waiver of **the republisher's own** rights over the collation. It
does not, and cannot, dispose of Survey of India's rights in the underlying
survey. Whether SoI's terms permitted this republication is a separate question
and nobody here has answered it.

Two things follow, and they point the same way. The residual risk is not
*our* compliance with CC0 — CC0 imposes nothing on us — but whether the grant was
the republisher's to make. And the same residual risk already applies to
`pc-boundaries.geojson`, which is CC0 over data DataMeet scraped from a
government website, and which this project has been serving without concern. So
this is not a reason to stop serving the file. It is a reason to name the
government source, which the credit line below does.

The republisher himself raises the point and quotes the government guideline that
bears on it — Indian Geospatial Guidelines 2021, item xiii: *"For political Maps
of India of any scale including national, state and other boundaries, SoI
published maps or SoI digital boundary data are the standard to be used, which
shall be made easily downloadable for free and their digital display and printing
shall be permissible."*

### Attribution text

There is no `attribution` member in this file to quote. **The following wording is
ours**, composed to honour the request quoted above:

> District boundaries: Survey of India, collated by the DataMeet community and
> republished by ramSeraph (indian_admin_boundaries), dedicated to the public
> domain under CC0 1.0.

It is held in `src/lib/geoAttribution.ts` and flagged there as composed rather
than transcribed. When the file is regenerated with an inline provenance block,
that block becomes the source of truth and this wording should be replaced by it.

### Dates

| | |
|---|---|
| Upstream release published | 2026-03-08 (GitHub API, `admin/districts`) |
| Committed to this repository | 2026-03-27, in `6119ac8` |
| **Therefore fetched** | between 2026-03-08 and 2026-03-27 |
| Geometry vintage, upper bound | on or before 2023-12-11, the origin release's publication date |
| Aggregator's own `last_updated` | 2024-08-15 |

The 2026-03-27 filesystem mtime is not evidence of anything; the commit date is.

**Geometry vintage — our reading from the file's own contents**, labelled as ours
because no provenance record states it. Ladakh appears as a separate `state_name`
and Dadra & Nagar Haveli and Daman & Diu as a single merged one, so the layer
postdates those reorganisations. Andhra Pradesh appears with 13 districts, which
is its pre-2022 count, so the layer predates that reorganisation. Individual
district creations inside that span are not reflected consistently, so do not
treat the layer as current for any particular district.

### Known defects

1. **Eight polygons are not districts.** Their `state_name` is a disputed-area
   label: `DISPUTED (MADHYA PRADESH & RAJASTHAN)` ×4,
   `DISPUTED (RAJATHAN & GUJARAT)` ×2 — the typo is the source's —
   `DISPUTED (MADHYA PRADESH & GUJARAT)`, and
   `DISPUTED (WEST BENGAL , BIHAR & JHARKHAND)`. They receive `district_id`s like
   `disputed_madhya_pradesh_rajasthan_disputed_ratlam_banswara` and will be
   treated as districts by any consumer that does not filter them. 742 features
   are therefore 734 districts plus 8 disputed areas.
2. **40 distinct `state_name` values**, which is not a count of states and UTs, for
   the reason above.
3. **Names carry source formatting.** `DAKSHINA  KANNADA` and `UTTARA  KANNADA`
   contain a doubled space. Normalise on read; do not edit the file in place
   without re-recording provenance.
4. **The shipped file carries no provenance whatsoever** — its only top-level keys
   are `type` and `features`, unlike the two constituency layers, which carry
   `license`, `attribution`, `source_url`, `source_sha256`, `fetched` and
   `known_defects` inline. Until `scripts/prepare-districts.mjs` writes that block
   and the file is regenerated, the licence does not travel with the bytes and a
   copy of this file has nothing attached to it.
5. **742 features have never been reconciled** against the Local Government
   Directory district count for any date. The layer's currency is a reading, not a
   record.

### Public export

**Yes**, with the credit line above carried. CC0 imposes no condition. This
supersedes the earlier position that the file was blocked pending a licence
answer: the answer is CC0, sourced above.

| | |
|---|---|
| Our file SHA-256 | `3e9700892757c3b06f60e1dfa1d421447656f4f67e0b6651c0b69cac5b4e381a` |
| Served at | `https://interns.city/data/india-districts.geojson` (HTTP 200, 3,153,362 bytes, 2026-09-30) |

---

## 4. `district-meta.json`

A name index derived from `india-districts.geojson`: 742 keys, each
`district_id → { district_name, state_name }`. No geometry.

Same licence, same source, same credit line and the same defects as §3, because
it is that file with the coordinates dropped. SHA-256
`34324af949d21283a646767a00ab4e26fc833d03040f6f937444409b454de19d`.

**It used to carry `population` and `area_sq_km`.** Those were computed as
`stateTotal / districtCount * Math.random()` and were wrong by up to roughly 180×
on real districts. Both fields and the UI that displayed them were removed. Do not
re-add them without a real per-district source — Census 2011 district primary
census abstract, or MoSPI — wired in as a required input.

---

## Rules that keep this file true

- **Record the licence per file, never per repository.** Every wrong statement
  this document had to unpick was a repository-level badge or a detector's guess
  being read as a claim about one file. `datameet/maps` carries three different
  licences in three directory READMEs. `india-geodata` shows a CC BY 4.0 badge
  over a CC0 release. Badges are not licences.
- **Read down to the most specific statement, then stop.** Repository badge, then
  `LICENSE`, then the dataset's `metadata.json`, then the directory README, then
  the release note for the individual asset. The answer was in the last two.
- **Put the licence inside the data file.** Two of the three geojson files here
  carry their licence as collection-level members, so the obligation cannot be
  separated from the bytes by a copy. The third does not, and §3 is the cost.
- **A file in `public/` is published.** Not internal, not "just for the resolver".
  An attribution obligation is live from the moment it deploys.
- **Distinguish a waiver from a grant.** CC0 from a republisher relinquishes that
  republisher's rights. It says nothing about the government source underneath.
  Name the government source.
