Compare commits

..

3 Commits

Author SHA1 Message Date
rosemariejebbjtxbfp 19b7c43f37 ci: connect to postgres via service hostname for Gitea Actions
CI / Go (api) (pull_request) Successful in 30s
CI / Python (ingestion) (pull_request) Successful in 7s
CI / Migrations (postgres) (pull_request) Successful in 16s
Gitea Actions runs jobs inside a container, so a service container is reachable by its service name (postgres), not localhost; localhost:5432 yields connection refused. Also drop the unnecessary 5432:5432 host port mapping that caused 'port already allocated' collisions.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-19 08:54:18 +00:00
lixu 57537c880d M4: ingestion management (#1)
CI / Go (api) (push) Failing after 14s
CI / Migrations (postgres) (push) Failing after 33s
CI / Python (ingestion) (push) Failing after 13m10s
Incremental OFF + GS1 supplement + dedup/conflict + quality scoring + scheduler
2026-06-08 17:40:26 +08:00
John Doe a35bcd6647 M4: ingestion management (incremental, GS1 supplement, dedup/conflict, quality, scheduler)
CI / Go (api) (pull_request) Has been cancelled
CI / Python (ingestion) (pull_request) Has been cancelled
CI / Migrations (postgres) (pull_request) Has been cancelled
- OFF incremental fetch via search API + persistent watermark (ingest_state, migration 0004)
- GS1 barcode supplement adapter (offline mapping + GS1-style API) filling only gaps with field-level provenance
- Non-GTIN dedup with canonical selection + merge_log; field-level conflict resolution (source trust > recency)
- Quality scoring (0.4 completeness + 0.3 source trust + 0.2 multi-source + 0.1 freshness) wired into load/merge
- Jobs: update_off, dedup, schedule; docs/ingestion-management.md
- 19 new tests (pure + DB-integration), ruff clean

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-08 09:27:42 +00:00
24 changed files with 1389 additions and 9 deletions
+2 -6
View File
@@ -19,13 +19,11 @@ jobs:
POSTGRES_USER: opengoods
POSTGRES_PASSWORD: opengoods
POSTGRES_DB: opengoods
ports:
- "5432:5432"
options: >-
--health-cmd "pg_isready -U opengoods"
--health-interval 5s --health-timeout 5s --health-retries 10
env:
OPENGOODS_DATABASE_URL: postgres://opengoods:opengoods@localhost:5432/opengoods?sslmode=disable
OPENGOODS_DATABASE_URL: postgres://opengoods:opengoods@postgres:5432/opengoods?sslmode=disable
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
@@ -73,13 +71,11 @@ jobs:
POSTGRES_USER: opengoods
POSTGRES_PASSWORD: opengoods
POSTGRES_DB: opengoods
ports:
- "5432:5432"
options: >-
--health-cmd "pg_isready -U opengoods"
--health-interval 5s --health-timeout 5s --health-retries 10
env:
DBURL: postgres://opengoods:opengoods@localhost:5432/opengoods?sslmode=disable
DBURL: postgres://opengoods:opengoods@postgres:5432/opengoods?sslmode=disable
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
+76
View File
@@ -0,0 +1,76 @@
# 采集管理 (M4)
M4 在 M2Open Food Facts 首次导入)基础上,补齐"持续运营"所需的采集能力:
增量更新、第二数据源补全(GS1)、去重合并与字段级冲突解决、数据质量评分,
以及把这些串起来的定时调度。全部为 Python 侧(`ingestion/`),只写库、可单测。
## 组成
| 能力 | 模块 | 说明 |
|------|------|------|
| 增量采集 | `adapters/openfoodfacts.py: fetch_modified_since()` | 按 `last_modified_t` 拉取自上次水位后变更的商品 |
| 采集水位 | `etl/state.py` + `ingest_state` 表 | 每个源持久化 `last_modified_t`,只前进不回退 |
| GS1 补全 | `adapters/gs1.py` + `etl/supplement.py` | 用权威条码源补**缺失**字段(品牌/厂商/GPC/产地/净含量),不覆盖已有值 |
| 去重合并 | `etl/dedup.py` | 非 GTIN 重复(同名+品牌+净含量)合并到质量最高的主记录 |
| 冲突解决 | `etl/merge.py` | 多源同字段按"源权重 > 新鲜度"择优,保留字段级溯源 |
| 质量评分 | `etl/quality.py` | 0~1 分,落到 `product.quality_score` |
| 定时调度 | `jobs/schedule.py` | 固定周期跑"增量 + 去重"一轮,零额外依赖 |
## 质量评分
锁定公式(各分量均归一到 0~1):
```
quality = 0.4 * 完整度 + 0.3 * 源权重 + 0.2 * 多源一致 + 0.1 * 新鲜度
```
- **完整度**`name/gtin/brand/category/net_content/country/nutriments/ingredients/image` 9 项的命中比例。
- **源权重**:贡献该商品的源中最高 `source.trust_weight`OFF=0.7GS1=0.9)。
- **多源一致**:源数量代理——单源 0.5、两源 0.8、三源及以上 1.0(单源无法互证)。
- **新鲜度**:最近一次 `product_source.fetched_at` 的时间衰减(≤30d=1.0 … >730d=0.2)。
`load_record()``merge_products()` 写入后都会调 `update_quality()` 重算。
## 增量水位
`ingest_state`(迁移 `0004`)每源一行,记录 `last_modified_t``last_run_at``stats`
`set_watermark()``GREATEST(...)` 保证水位只前进,避免乱序/中断的运行回退进度。
## 运行
前置:`docker compose up -d postgres` 且迁移已 `up`(含 `0004`)。DSN 默认读 `OPENGOODS_DATABASE_URL`
```bash
# 增量更新 OFF(从持久化水位开始;--since 可覆盖)
python -m opengoods.jobs.update_off --max-pages 5
python -m opengoods.jobs.update_off --since 1700000000
# 去重合并(--dry-run 只报告不写库)
python -m opengoods.jobs.dedup --dry-run
python -m opengoods.jobs.dedup --actor nightly
# 定时调度:单轮 / 周期循环(增量 + 去重)
python -m opengoods.jobs.schedule --once
python -m opengoods.jobs.schedule --interval 3600
```
## GS1 补全
GS1 为付费、分区域的授权数据,适配器支持两种模式:
- **离线**(默认):从本地 JSON 映射 `{gtin: {...}}` 查(`GS1Adapter.from_file(path)`),
供测试与内网环境使用。
- **在线**:传 `base_url` + `client`+ `api_key`),`GET {base_url}/{gtin}`,按
Verified-by-GS1 风格字段解析。
补全只填**空缺**字段并在 `product_source` 记字段级溯源,源标记为 `gs1`
## 测试
```bash
cd ingestion && pip install -e ".[dev]"
ruff check . && ruff format --check . && pytest -q
```
纯函数测试(质量/冲突/增量分页)始终运行;依赖库的测试(水位/质量落库/去重/GS1 补全)
在无数据库或未应用 M4 迁移时自动跳过。
+124
View File
@@ -0,0 +1,124 @@
"""GS1 barcode supplement adapter.
GS1 (e.g. *Verified by GS1* / GS1 China) is the authoritative registry that maps
a GTIN to its brand owner, product description and GPC category. We use it to
*supplement* — fill gaps in — records gathered from crowd sources like Open Food
Facts, never to overwrite existing values.
Real GS1 access is credentialed and region-specific, so this adapter supports
two modes:
* **offline** (default): look barcodes up in a local JSON mapping file. This is
what tests and air-gapped runs use.
* **online**: GET ``{base_url}/{gtin}`` with an API key header, then normalize
the response. Enabled by passing ``base_url`` + ``client``.
Either way :meth:`fetch_barcode` returns a normalized *supplement* dict (or
``None``); :mod:`opengoods.etl.supplement` applies it to the database.
"""
from __future__ import annotations
import json
from collections.abc import Iterator
from pathlib import Path
import httpx
SOURCE_NAME = "gs1"
GS1_HOMEPAGE = "https://www.gs1.org"
GS1_LICENSE = "proprietary"
# GS1 is the authoritative barcode registry -> high trust.
GS1_TRUST = 0.9
# Keys of a normalized supplement record.
_SUPPLEMENT_KEYS = (
"gtin",
"name",
"brand",
"manufacturer",
"gpc_brick_code",
"country_of_origin",
"net_content_value",
"net_content_unit",
)
def _normalize(code: str, data: dict) -> dict:
"""Project a raw mapping/record onto the supplement schema (non-empty only)."""
rec: dict = {"gtin": code}
for key in _SUPPLEMENT_KEYS:
if key == "gtin":
continue
value = data.get(key)
if value not in (None, "", []):
rec[key] = value
return rec
def _parse_api(code: str, payload: dict) -> dict:
"""Best-effort mapping of a Verified-by-GS1 style payload to our schema."""
item = payload
if isinstance(payload.get("gtinRecords"), list) and payload["gtinRecords"]:
item = payload["gtinRecords"][0]
return _normalize(
code,
{
"name": item.get("productDescription") or item.get("description"),
"brand": item.get("brandName"),
"manufacturer": item.get("companyName") or item.get("licenseeName"),
"gpc_brick_code": item.get("gpcCategoryCode"),
"country_of_origin": item.get("countryOfSaleCode") or item.get("countryCode"),
"net_content_value": item.get("netContent"),
"net_content_unit": item.get("netContentUnit"),
},
)
class GS1Adapter:
"""Look up GTIN supplements from a local mapping or a GS1-style API."""
source_name = SOURCE_NAME
def __init__(
self,
mapping: dict | None = None,
*,
client: httpx.Client | None = None,
base_url: str | None = None,
api_key: str | None = None,
) -> None:
self._mapping = mapping or {}
self._client = client
self._base_url = base_url.rstrip("/") if base_url else None
self._api_key = api_key
@classmethod
def from_file(cls, path: str | Path) -> GS1Adapter:
"""Build an offline adapter from a JSON ``{gtin: {...}}`` mapping file."""
data = json.loads(Path(path).read_text(encoding="utf-8"))
return cls(mapping=data)
def fetch_barcode(self, code: str) -> dict | None:
"""Return a normalized supplement dict for ``code`` (or ``None``)."""
if self._base_url and self._client is not None:
headers = {"apikey": self._api_key} if self._api_key else {}
resp = self._client.get(f"{self._base_url}/{code}", headers=headers)
if resp.status_code == 404:
return None
resp.raise_for_status()
rec = _parse_api(code, resp.json())
else:
data = self._mapping.get(code)
if not data:
return None
rec = _normalize(code, data)
# A record with only the GTIN carries no supplement.
return rec if len(rec) > 1 else None
def fetch(self, barcodes: list[str]) -> Iterator[dict]:
"""Yield supplement records for the given barcodes."""
for code in barcodes:
rec = self.fetch_barcode(code)
if rec is not None:
yield rec
@@ -25,6 +25,16 @@ USER_AGENT = "OpenGoods/0.1 (+https://github.com/baicai2026-baicai/goods) public
# Conservative client-side spacing between API calls (seconds).
_DEFAULT_MIN_INTERVAL = 4.0
_API_URL = "https://world.openfoodfacts.org/api/v2/product/{barcode}.json"
_SEARCH_URL = "https://world.openfoodfacts.org/api/v2/search"
# Fields requested from the search API so a returned product can be transformed
# without an extra per-barcode round trip.
_SEARCH_FIELDS = (
"code,product_name,product_name_en,product_name_zh,brands,quantity,"
"categories,categories_tags,countries,ingredients_text,allergens_tags,"
"additives_tags,nutriments,nutriscore_grade,serving_size,"
"image_front_url,image_url,last_modified_t"
)
class OpenFoodFactsAdapter:
@@ -65,6 +75,45 @@ class OpenFoodFactsAdapter:
if record is not None:
yield record
def fetch_modified_since(
self,
since_t: int,
*,
page_size: int = 100,
max_pages: int = 10,
) -> Iterator[dict]:
"""Yield products modified after ``since_t`` (unix ``last_modified_t``).
Uses the OFF search API sorted by ``last_modified_t`` (most recent
first) and paginates until it reaches products at or before the
watermark, an empty/short page, or ``max_pages``. This is the
incremental ingestion path: callers persist the highest
``last_modified_t`` they processed as the next watermark.
"""
for page in range(1, max_pages + 1):
self._throttle()
resp = self._client.get(
_SEARCH_URL,
params={
"fields": _SEARCH_FIELDS,
"sort_by": "last_modified_t",
"page": page,
"page_size": page_size,
},
)
resp.raise_for_status()
products = resp.json().get("products") or []
if not products:
return
reached_old = False
for prod in products:
if int(prod.get("last_modified_t") or 0) <= since_t:
reached_old = True
break
yield prod
if reached_old or len(products) < page_size:
return
def read_dump(path: str | Path) -> Iterator[dict]:
"""Yield raw product records from an OFF JSONL dump file.
+139
View File
@@ -0,0 +1,139 @@
"""Duplicate detection and product merging.
Barcodes (GTIN) are already unique at the schema level, so duplicates here are
non-GTIN records that describe the same product (same normalized name + brand +
net content). For each duplicate group we keep the highest-quality product as
canonical and merge the rest into it: child rows (provenance, images, MSRP) are
re-pointed to the canonical product, the merged product is marked ``merged``
with ``canonical_id`` set, and a row is written to ``merge_log``.
"""
from __future__ import annotations
from typing import Any
import psycopg
from opengoods.etl.quality import update_quality
def _norm(text: str | None) -> str:
return " ".join((text or "").lower().split())
def product_signature(name: str | None, brand: str | None, net_canonical: Any | None) -> str | None:
"""Stable signature for non-GTIN dedup, or ``None`` if too sparse to match."""
n = _norm(name)
if not n:
return None
net = "" if net_canonical is None else str(net_canonical)
return f"{n}|{_norm(brand)}|{net}"
def choose_canonical(members: list[dict]) -> dict:
"""Pick the canonical product: best quality, then oldest, then lowest id."""
return min(
members,
key=lambda m: (
-float(m.get("quality_score") or 0.0),
m.get("created_at"),
str(m.get("id")),
),
)
def find_duplicate_groups(conn: psycopg.Connection) -> list[list[dict]]:
"""Return groups (size >= 2) of active products sharing a signature."""
rows = conn.execute(
"""
SELECT p.id, p.name, b.normalized_name, p.net_content_canonical,
p.quality_score, p.created_at
FROM product p
LEFT JOIN brand b ON b.id = p.brand_id
WHERE p.status = 'active'
"""
).fetchall()
groups: dict[str, list[dict]] = {}
for r in rows:
sig = product_signature(r[1], r[2], r[3])
if sig is None:
continue
member = {
"id": r[0],
"name": r[1],
"quality_score": r[4],
"created_at": r[5],
}
groups.setdefault(sig, []).append(member)
return [m for m in groups.values() if len(m) >= 2]
def merge_products(
conn: psycopg.Connection,
kept_id: str,
merged_id: str,
reason: str = "auto-dedup",
actor: str = "ingestion",
) -> None:
"""Merge ``merged_id`` into ``kept_id`` (re-point children, mark merged)."""
if kept_id == merged_id:
return
# Re-point provenance, images and MSRP to the canonical product.
conn.execute(
"UPDATE product_source SET product_id = %s WHERE product_id = %s",
(kept_id, merged_id),
)
conn.execute(
"UPDATE product_image SET product_id = %s WHERE product_id = %s",
(kept_id, merged_id),
)
conn.execute(
"UPDATE product_msrp SET product_id = %s WHERE product_id = %s",
(kept_id, merged_id),
)
# food_detail has product_id as PK, so it can only move if the canonical
# product does not already have one.
kept_has_food = conn.execute(
"SELECT 1 FROM food_detail WHERE product_id = %s", (kept_id,)
).fetchone()
if not kept_has_food:
conn.execute(
"UPDATE food_detail SET product_id = %s WHERE product_id = %s",
(kept_id, merged_id),
)
conn.execute(
"UPDATE product SET status = 'merged', canonical_id = %s WHERE id = %s",
(kept_id, merged_id),
)
conn.execute(
"""
INSERT INTO merge_log (kept_id, merged_id, reason, actor)
VALUES (%s, %s, %s, %s)
""",
(kept_id, merged_id, reason, actor),
)
# The canonical product gained sources, so its quality may have changed.
update_quality(conn, kept_id)
def dedup_all(
conn: psycopg.Connection, actor: str = "ingestion", dry_run: bool = False
) -> dict[str, int]:
"""Merge every duplicate group. Returns counts of groups and merges."""
groups = find_duplicate_groups(conn)
merged = 0
for members in groups:
canonical = choose_canonical(members)
for m in members:
if m["id"] == canonical["id"]:
continue
if not dry_run:
merge_products(conn, canonical["id"], m["id"], actor=actor)
merged += 1
return {"groups": len(groups), "merged": merged}
+18 -3
View File
@@ -14,6 +14,7 @@ import psycopg
from psycopg.types.json import Jsonb
from opengoods.adapters.openfoodfacts import OFF_LICENSE, SOURCE_NAME
from opengoods.etl.quality import update_quality
OFF_HOMEPAGE = "https://world.openfoodfacts.org"
@@ -29,8 +30,14 @@ def _normalize_brand(name: str) -> str:
return " ".join(name.lower().split())
def ensure_source(conn: psycopg.Connection) -> str:
"""Upsert the Open Food Facts source row and return its id."""
def ensure_source_named(
conn: psycopg.Connection,
name: str,
homepage: str,
license: str,
trust_weight: float,
) -> str:
"""Upsert a source row by name and return its id."""
row = conn.execute(
"""
INSERT INTO source (name, homepage, license, trust_weight)
@@ -38,11 +45,16 @@ def ensure_source(conn: psycopg.Connection) -> str:
ON CONFLICT (name) DO UPDATE SET homepage = EXCLUDED.homepage
RETURNING id
""",
(SOURCE_NAME, OFF_HOMEPAGE, OFF_LICENSE, 0.7),
(name, homepage, license, trust_weight),
).fetchone()
return row[0]
def ensure_source(conn: psycopg.Connection) -> str:
"""Upsert the Open Food Facts source row and return its id."""
return ensure_source_named(conn, SOURCE_NAME, OFF_HOMEPAGE, OFF_LICENSE, 0.7)
def _ensure_brand(conn: psycopg.Connection, name: str | None) -> str | None:
if not name:
return None
@@ -178,6 +190,9 @@ def load_record(conn: psycopg.Connection, rec: dict[str, Any], source_id: str, r
Jsonb(_jsonable(raw)),
),
)
# Recompute the data-quality score now that all facts + provenance exist.
update_quality(conn, product_id)
return product_id
+104
View File
@@ -0,0 +1,104 @@
"""Field-level conflict resolution for multi-source records.
When more than one source describes the same product, each field may have
several candidate values. We pick a winner per field by source trust first,
then recency, ignoring empty values, and keep a provenance trail of which
source won each field.
These are pure functions (no DB / no network) so they are easy to unit-test;
the DB-level record merge lives in :mod:`opengoods.etl.dedup`.
"""
from __future__ import annotations
from dataclasses import dataclass, field
from datetime import datetime
@dataclass(frozen=True)
class Candidate:
"""One source's proposed value for a field."""
value: object
source: str
trust: float = 0.5
fetched_at: datetime | None = None
@dataclass
class FieldResolution:
"""The winning value for a field plus the source it came from."""
value: object
source: str | None = None
@dataclass
class MergedRecord:
"""A merged record with per-field provenance (field name -> source)."""
values: dict[str, object] = field(default_factory=dict)
provenance: dict[str, str] = field(default_factory=dict)
def _is_empty(value: object) -> bool:
if value is None:
return True
if isinstance(value, str):
return value.strip() == ""
if isinstance(value, (list, dict, tuple, set)):
return len(value) == 0
return False
def _sort_key(c: Candidate) -> tuple[float, float]:
ts = c.fetched_at.timestamp() if c.fetched_at is not None else float("-inf")
return (c.trust, ts)
def resolve_field(candidates: list[Candidate]) -> FieldResolution | None:
"""Pick the best non-empty candidate for one field.
Ranking: highest source trust, then most recent ``fetched_at``. Returns
``None`` when there is no usable (non-empty) candidate.
"""
usable = [c for c in candidates if not _is_empty(c.value)]
if not usable:
return None
winner = max(usable, key=_sort_key)
return FieldResolution(value=winner.value, source=winner.source)
def merge_records(records: list[dict], *, fields: list[str] | None = None) -> MergedRecord:
"""Merge several ``{field: Candidate|value}`` records into one.
Each input record maps field name -> :class:`Candidate` (preferred) or a
bare value (treated as trust 0.5, no timestamp). The result keeps, for each
field, the winning value and the name of the source that supplied it.
"""
keys: list[str]
if fields is not None:
keys = list(fields)
else:
seen: dict[str, None] = {}
for rec in records:
for k in rec:
seen.setdefault(k, None)
keys = list(seen)
merged = MergedRecord()
for key in keys:
candidates: list[Candidate] = []
for rec in records:
if key not in rec:
continue
cand = rec[key]
if not isinstance(cand, Candidate):
cand = Candidate(value=cand, source="unknown")
candidates.append(cand)
resolution = resolve_field(candidates)
if resolution is not None:
merged.values[key] = resolution.value
if resolution.source is not None:
merged.provenance[key] = resolution.source
return merged
+176
View File
@@ -0,0 +1,176 @@
"""Product data-quality scoring.
The quality score is a 0..1 number combining four signals, per the locked
project decision:
quality = 0.4 * completeness
+ 0.3 * source_trust
+ 0.2 * multi_source_agreement
+ 0.1 * freshness
Each component is itself normalized to 0..1. The pure helpers below are
unit-testable; :func:`compute_quality` / :func:`update_quality` read the signals
for a product out of the database and persist the result on ``product``.
"""
from __future__ import annotations
from datetime import UTC, datetime
import psycopg
W_COMPLETENESS = 0.4
W_SOURCE_TRUST = 0.3
W_AGREEMENT = 0.2
W_FRESHNESS = 0.1
# Fields that count towards completeness (weighted equally).
COMPLETENESS_FIELDS = (
"name",
"gtin",
"brand",
"category",
"net_content",
"country_of_origin",
"nutriments",
"ingredients",
"image",
)
def completeness(present: set[str]) -> float:
"""Fraction of :data:`COMPLETENESS_FIELDS` that are present for a product."""
if not COMPLETENESS_FIELDS:
return 0.0
hits = sum(1 for f in COMPLETENESS_FIELDS if f in present)
return hits / len(COMPLETENESS_FIELDS)
def agreement_from_sources(source_count: int) -> float:
"""Multi-source corroboration proxy from the number of distinct sources.
A single source cannot be corroborated, so it scores a neutral 0.5; more
independent sources that describe the same product raise confidence.
"""
if source_count <= 1:
return 0.5
if source_count == 2:
return 0.8
return 1.0
def freshness_from_age(age_days: float | None) -> float:
"""Recency score from the age (in days) of the most recent source fetch."""
if age_days is None:
return 0.5
if age_days <= 30:
return 1.0
if age_days <= 180:
return 0.8
if age_days <= 365:
return 0.6
if age_days <= 730:
return 0.4
return 0.2
def score(
*,
completeness_score: float,
source_trust: float,
agreement: float,
freshness: float,
) -> float:
"""Combine the four normalized components into a 0..1 quality score."""
raw = (
W_COMPLETENESS * completeness_score
+ W_SOURCE_TRUST * source_trust
+ W_AGREEMENT * agreement
+ W_FRESHNESS * freshness
)
return round(max(0.0, min(1.0, raw)), 3)
def _present_fields(prod: dict, has_image: bool) -> set[str]:
present: set[str] = set()
if prod.get("name"):
present.add("name")
if prod.get("gtin"):
present.add("gtin")
if prod.get("brand_id"):
present.add("brand")
if prod.get("category_id"):
present.add("category")
if prod.get("net_content_canonical") is not None:
present.add("net_content")
if prod.get("country_of_origin"):
present.add("country_of_origin")
if prod.get("nutriments"):
present.add("nutriments")
if prod.get("ingredients_text"):
present.add("ingredients")
if has_image:
present.add("image")
return present
def compute_quality(conn: psycopg.Connection, product_id: str) -> float:
"""Compute (but do not persist) the quality score for one product."""
row = conn.execute(
"""
SELECT p.name, p.gtin, p.brand_id, p.category_id, p.net_content_canonical,
p.country_of_origin, f.nutriments, f.ingredients_text,
EXISTS (SELECT 1 FROM product_image pi WHERE pi.product_id = p.id)
FROM product p
LEFT JOIN food_detail f ON f.product_id = p.id
WHERE p.id = %s
""",
(product_id,),
).fetchone()
if row is None:
return 0.0
prod = {
"name": row[0],
"gtin": row[1],
"brand_id": row[2],
"category_id": row[3],
"net_content_canonical": row[4],
"country_of_origin": row[5],
"nutriments": row[6],
"ingredients_text": row[7],
}
has_image = bool(row[8])
src = conn.execute(
"""
SELECT count(DISTINCT ps.source_id), COALESCE(max(s.trust_weight), 0), max(ps.fetched_at)
FROM product_source ps
LEFT JOIN source s ON s.id = ps.source_id
WHERE ps.product_id = %s
""",
(product_id,),
).fetchone()
source_count = int(src[0] or 0)
source_trust = float(src[1] or 0.0)
last_fetched: datetime | None = src[2]
age_days: float | None = None
if last_fetched is not None:
now = datetime.now(UTC)
if last_fetched.tzinfo is None:
last_fetched = last_fetched.replace(tzinfo=UTC)
age_days = max(0.0, (now - last_fetched).total_seconds() / 86400.0)
return score(
completeness_score=completeness(_present_fields(prod, has_image)),
source_trust=source_trust,
agreement=agreement_from_sources(source_count),
freshness=freshness_from_age(age_days),
)
def update_quality(conn: psycopg.Connection, product_id: str) -> float:
"""Compute the quality score and write it to ``product.quality_score``."""
value = compute_quality(conn, product_id)
conn.execute("UPDATE product SET quality_score = %s WHERE id = %s", (value, product_id))
return value
+45
View File
@@ -0,0 +1,45 @@
"""Persistent ingestion watermark stored in the ``ingest_state`` table.
The incremental updater uses this to remember how far it got for each source
(e.g. Open Food Facts exposes a ``last_modified_t`` unix timestamp on every
product) so repeated runs only fetch what changed.
"""
from __future__ import annotations
from typing import Any
import psycopg
from psycopg.types.json import Jsonb
def get_watermark(conn: psycopg.Connection, source: str) -> int:
"""Return the last processed ``last_modified_t`` for *source* (0 if none)."""
row = conn.execute(
"SELECT last_modified_t FROM ingest_state WHERE source = %s", (source,)
).fetchone()
return int(row[0]) if row else 0
def set_watermark(
conn: psycopg.Connection,
source: str,
last_modified_t: int,
stats: dict[str, Any] | None = None,
) -> None:
"""Upsert the watermark and run metadata for *source*.
The watermark only ever moves forward: a lower ``last_modified_t`` is
ignored so an out-of-order or partial run cannot rewind progress.
"""
conn.execute(
"""
INSERT INTO ingest_state (source, last_modified_t, last_run_at, stats)
VALUES (%s, %s, now(), %s)
ON CONFLICT (source) DO UPDATE SET
last_modified_t = GREATEST(ingest_state.last_modified_t, EXCLUDED.last_modified_t),
last_run_at = now(),
stats = EXCLUDED.stats
""",
(source, int(last_modified_t), Jsonb(stats or {})),
)
+133
View File
@@ -0,0 +1,133 @@
"""Apply GS1 (or other authoritative) supplements to existing products.
A supplement only fills *gaps*: a field is written only when the product does
not already have a value. Each applied supplement records field-level provenance
in ``product_source`` and refreshes the product's quality score.
"""
from __future__ import annotations
from decimal import Decimal, InvalidOperation
from typing import Any
import psycopg
from psycopg.types.json import Jsonb
from opengoods import units
from opengoods.adapters.gs1 import GS1_HOMEPAGE, GS1_LICENSE, GS1_TRUST, SOURCE_NAME
from opengoods.etl.load import _ensure_brand, _normalize_brand, ensure_source_named
from opengoods.etl.quality import update_quality
def ensure_gs1_source(conn: psycopg.Connection) -> str:
"""Upsert the GS1 source row and return its id."""
return ensure_source_named(conn, SOURCE_NAME, GS1_HOMEPAGE, GS1_LICENSE, GS1_TRUST)
def _ensure_manufacturer(conn: psycopg.Connection, name: str | None) -> str | None:
if not name:
return None
row = conn.execute(
"""
INSERT INTO manufacturer (name, normalized_name)
VALUES (%s, %s)
ON CONFLICT (normalized_name) DO UPDATE SET name = manufacturer.name
RETURNING id
""",
(name, _normalize_brand(name)),
).fetchone()
return row[0]
def _net_content(rec: dict) -> tuple[Decimal, str, Decimal | None] | None:
raw_value = rec.get("net_content_value")
unit = rec.get("net_content_unit")
if raw_value is None or not unit:
return None
try:
value = Decimal(str(raw_value))
except (InvalidOperation, ValueError):
return None
try:
canonical = units.normalize(value, unit).canonical_value
except units.UnitError:
canonical = None
return value, unit, canonical
def apply_supplement(conn: psycopg.Connection, rec: dict[str, Any], source_id: str) -> list[str]:
"""Fill missing fields of the GTIN-matched product from ``rec``.
Returns the list of field names actually filled (empty if the product is
unknown or already complete for the supplied fields).
"""
gtin = rec.get("gtin")
if not gtin:
return []
prod = conn.execute(
"""
SELECT id, brand_id, manufacturer_id, gpc_brick_code, country_of_origin,
net_content_value
FROM product
WHERE gtin = %s AND status = 'active'
""",
(gtin,),
).fetchone()
if prod is None:
return []
product_id, brand_id, manufacturer_id, gpc, country, net_value = prod
sets: list[str] = []
params: list[Any] = []
filled: list[str] = []
if brand_id is None and rec.get("brand"):
new_brand_id = _ensure_brand(conn, rec["brand"])
if new_brand_id is not None:
sets.append("brand_id = %s")
params.append(new_brand_id)
filled.append("brand")
if manufacturer_id is None and rec.get("manufacturer"):
new_mfr_id = _ensure_manufacturer(conn, rec["manufacturer"])
if new_mfr_id is not None:
sets.append("manufacturer_id = %s")
params.append(new_mfr_id)
filled.append("manufacturer")
if gpc is None and rec.get("gpc_brick_code"):
sets.append("gpc_brick_code = %s")
params.append(rec["gpc_brick_code"])
filled.append("gpc_brick_code")
if country is None and rec.get("country_of_origin"):
sets.append("country_of_origin = %s")
params.append(rec["country_of_origin"])
filled.append("country_of_origin")
if net_value is None:
net = _net_content(rec)
if net is not None:
value, unit, canonical = net
sets += [
"net_content_value = %s",
"net_content_unit = %s",
"net_content_canonical = %s",
]
params += [value, unit, canonical]
filled.append("net_content")
if not filled:
return []
params.append(product_id)
conn.execute(f"UPDATE product SET {', '.join(sets)} WHERE id = %s", params)
conn.execute(
"""
INSERT INTO product_source (product_id, source_id, url, fields, fetched_at, raw)
VALUES (%s, %s, %s, %s, now(), %s)
""",
(product_id, source_id, GS1_HOMEPAGE, filled, Jsonb(rec)),
)
update_quality(conn, product_id)
return filled
+40
View File
@@ -0,0 +1,40 @@
"""Deduplicate products: merge non-GTIN duplicates into a canonical record.
Usage:
python -m opengoods.jobs.dedup --dry-run
python -m opengoods.jobs.dedup --actor nightly
"""
from __future__ import annotations
import argparse
import sys
import psycopg
from opengoods.etl.dedup import dedup_all
from opengoods.etl.load import default_dsn
def run(args: argparse.Namespace) -> int:
with psycopg.connect(args.dsn, autocommit=False) as conn:
summary = dedup_all(conn, actor=args.actor, dry_run=args.dry_run)
if args.dry_run:
conn.rollback()
else:
conn.commit()
mode = "dry-run" if args.dry_run else "applied"
print(f"{mode} groups={summary['groups']} merged={summary['merged']}")
return 0
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Deduplicate OpenGoods products")
parser.add_argument("--actor", default="ingestion", help="merge_log actor label")
parser.add_argument("--dry-run", action="store_true", help="report only, do not write")
parser.add_argument("--dsn", default=default_dsn(), help="PostgreSQL DSN")
return run(parser.parse_args(argv))
if __name__ == "__main__":
sys.exit(main())
+66
View File
@@ -0,0 +1,66 @@
"""Lightweight recurring ingestion scheduler.
Runs one ingestion cycle (incremental OFF update, then dedup) on a fixed
interval. Dependency-free: a plain sleep loop rather than a cron/APScheduler
dependency, so it is trivial to run in a container or under systemd/supervisor.
Usage:
python -m opengoods.jobs.schedule --once # single cycle, then exit
python -m opengoods.jobs.schedule --interval 3600 # every hour
"""
from __future__ import annotations
import argparse
import sys
import time
from datetime import UTC, datetime
from opengoods.etl.load import default_dsn
from opengoods.jobs import dedup as dedup_job
from opengoods.jobs import update_off as update_job
def _cycle(args: argparse.Namespace) -> None:
ts = datetime.now(UTC).isoformat(timespec="seconds")
print(f"[{ts}] cycle start")
update_job.run(
argparse.Namespace(
since=None,
page_size=args.page_size,
max_pages=args.max_pages,
min_interval=args.min_interval,
dsn=args.dsn,
)
)
if not args.skip_dedup:
dedup_job.run(argparse.Namespace(actor="scheduler", dry_run=False, dsn=args.dsn))
def run(args: argparse.Namespace) -> int:
_cycle(args)
if args.once:
return 0
while True:
time.sleep(args.interval)
try:
_cycle(args)
except Exception as exc: # noqa: BLE001 - keep the loop alive across failures
print(f"cycle error: {exc}", file=sys.stderr)
return 0
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Recurring OpenGoods ingestion")
parser.add_argument("--interval", type=int, default=3600, help="seconds between cycles")
parser.add_argument("--once", action="store_true", help="run a single cycle and exit")
parser.add_argument("--skip-dedup", action="store_true", help="run update only")
parser.add_argument("--page-size", type=int, default=100, help="search page size")
parser.add_argument("--max-pages", type=int, default=10, help="max pages to scan")
parser.add_argument("--min-interval", type=float, default=4.0, help="API throttle seconds")
parser.add_argument("--dsn", default=default_dsn(), help="PostgreSQL DSN")
return run(parser.parse_args(argv))
if __name__ == "__main__":
sys.exit(main())
+65
View File
@@ -0,0 +1,65 @@
"""Incremental Open Food Facts update.
Fetches products modified since the persisted watermark, loads them, then
advances the watermark to the newest ``last_modified_t`` processed so the next
run only sees what changed.
Usage:
python -m opengoods.jobs.update_off --max-pages 5
python -m opengoods.jobs.update_off --since 1700000000 # override watermark
"""
from __future__ import annotations
import argparse
import sys
import psycopg
from opengoods.adapters.openfoodfacts import SOURCE_NAME, OpenFoodFactsAdapter
from opengoods.etl.load import default_dsn, ensure_source, load_record
from opengoods.etl.state import get_watermark, set_watermark
from opengoods.etl.transform import transform
def run(args: argparse.Namespace) -> int:
adapter = OpenFoodFactsAdapter(min_interval=args.min_interval)
loaded = skipped = 0
high_watermark = 0
with psycopg.connect(args.dsn, autocommit=False) as conn:
source_id = ensure_source(conn)
since = args.since if args.since is not None else get_watermark(conn, SOURCE_NAME)
high_watermark = since
for raw in adapter.fetch_modified_since(
since, page_size=args.page_size, max_pages=args.max_pages
):
high_watermark = max(high_watermark, int(raw.get("last_modified_t") or 0))
rec = transform(raw)
if rec is None:
skipped += 1
continue
load_record(conn, rec, source_id, raw)
loaded += 1
set_watermark(
conn,
SOURCE_NAME,
high_watermark,
stats={"loaded": loaded, "skipped": skipped, "since": since},
)
conn.commit()
print(f"since={since} loaded={loaded} skipped={skipped} watermark={high_watermark}")
return 0
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description="Incremental OFF update")
parser.add_argument("--since", type=int, default=None, help="override watermark (unix ts)")
parser.add_argument("--page-size", type=int, default=100, help="search page size")
parser.add_argument("--max-pages", type=int, default=10, help="max pages to scan")
parser.add_argument("--min-interval", type=float, default=4.0, help="API throttle seconds")
parser.add_argument("--dsn", default=default_dsn(), help="PostgreSQL DSN")
return run(parser.parse_args(argv))
if __name__ == "__main__":
sys.exit(main())
+31
View File
@@ -0,0 +1,31 @@
"""Shared test fixtures.
`db_conn` yields a psycopg connection inside a transaction that is rolled back
after each test, so DB tests stay isolated and leave no residue. Tests are
skipped automatically when no database is reachable or M4 migrations are not
applied (e.g. local runs without docker).
"""
from __future__ import annotations
import psycopg
import pytest
from opengoods.etl.load import default_dsn
@pytest.fixture()
def db_conn():
try:
conn = psycopg.connect(default_dsn(), connect_timeout=3)
except psycopg.OperationalError as exc: # pragma: no cover - env dependent
pytest.skip(f"no database available: {exc}")
has_state = conn.execute("SELECT to_regclass('public.ingest_state') IS NOT NULL").fetchone()[0]
if not has_state:
conn.close()
pytest.skip("M4 migrations not applied")
try:
yield conn
finally:
conn.rollback()
conn.close()
+14
View File
@@ -0,0 +1,14 @@
{
"06901234567892": {
"name": "示例矿泉水 550ml",
"brand": "示例品牌",
"manufacturer": "示例饮品有限公司",
"gpc_brick_code": "10000224",
"country_of_origin": "China",
"net_content_value": 550,
"net_content_unit": "ml"
},
"00000000000000": {
"brand": ""
}
}
+62
View File
@@ -0,0 +1,62 @@
from opengoods.etl.dedup import choose_canonical, dedup_all, product_signature
from opengoods.etl.load import _ensure_brand, ensure_source
def test_product_signature_normalization():
a = product_signature(" Spring Water ", "Acme", 500)
b = product_signature("spring water", "acme", 500)
assert a == b
assert product_signature("", "x", 1) is None
def test_choose_canonical_prefers_quality():
members = [
{"id": "a", "quality_score": 0.2, "created_at": 1},
{"id": "b", "quality_score": 0.9, "created_at": 2},
]
assert choose_canonical(members)["id"] == "b"
def test_dedup_merges_duplicates(db_conn):
brand_id = _ensure_brand(db_conn, "DupBrand")
def mk(quality):
return db_conn.execute(
"""
INSERT INTO product (name, brand_id, net_content_canonical, quality_score)
VALUES (%s, %s, %s, %s) RETURNING id
""",
("Dup Snack", brand_id, 100, quality),
).fetchone()[0]
keep = mk(0.9)
drop = mk(0.2)
src = ensure_source(db_conn)
db_conn.execute(
"INSERT INTO product_source (product_id, source_id, fields) VALUES (%s, %s, %s)",
(drop, src, ["name"]),
)
summary = dedup_all(db_conn)
assert summary == {"groups": 1, "merged": 1}
keep_status = db_conn.execute("SELECT status FROM product WHERE id = %s", (keep,)).fetchone()[0]
drop_status, canonical_id = db_conn.execute(
"SELECT status, canonical_id FROM product WHERE id = %s", (drop,)
).fetchone()
assert keep_status == "active"
assert drop_status == "merged"
assert str(canonical_id) == str(keep)
# The merged product's source row was re-pointed to the canonical product.
reattached = db_conn.execute(
"SELECT count(*) FROM product_source WHERE product_id = %s", (keep,)
).fetchone()[0]
assert reattached == 1
logged = db_conn.execute(
"SELECT count(*) FROM merge_log WHERE kept_id = %s AND merged_id = %s",
(keep, drop),
).fetchone()[0]
assert logged == 1
+53
View File
@@ -0,0 +1,53 @@
from pathlib import Path
from opengoods.adapters.gs1 import GS1Adapter
from opengoods.etl.supplement import apply_supplement, ensure_gs1_source
MAPPING = Path(__file__).parent / "fixtures" / "gs1_mapping.json"
GTIN = "06901234567892"
def test_gs1_adapter_offline_lookup():
adapter = GS1Adapter.from_file(MAPPING)
rec = adapter.fetch_barcode(GTIN)
assert rec["brand"] == "示例品牌"
assert rec["net_content_value"] == 550
assert rec["net_content_unit"] == "ml"
# An entry that only has empty values yields no supplement.
assert adapter.fetch_barcode("00000000000000") is None
# Unknown barcode -> None.
assert adapter.fetch_barcode("99999999999999") is None
def test_gs1_supplement_fills_only_gaps(db_conn):
pid = db_conn.execute(
"INSERT INTO product (gtin, name) VALUES (%s, %s) RETURNING id", (GTIN, "")
).fetchone()[0]
adapter = GS1Adapter.from_file(MAPPING)
rec = adapter.fetch_barcode(GTIN)
source_id = ensure_gs1_source(db_conn)
filled = apply_supplement(db_conn, rec, source_id)
assert {"brand", "country_of_origin", "net_content"} <= set(filled)
brand_id, country, net_value, net_unit = db_conn.execute(
"""
SELECT brand_id, country_of_origin, net_content_value, net_content_unit
FROM product WHERE id = %s
""",
(pid,),
).fetchone()
assert brand_id is not None
assert country == "China"
assert float(net_value) == 550.0
assert net_unit == "ml"
fields = db_conn.execute(
"SELECT fields FROM product_source WHERE product_id = %s AND source_id = %s",
(pid, source_id),
).fetchone()[0]
assert "brand" in fields
# Re-applying does nothing because the gaps are now filled.
assert apply_supplement(db_conn, rec, source_id) == []
+62
View File
@@ -0,0 +1,62 @@
from datetime import UTC, datetime
from opengoods.etl.merge import Candidate, merge_records, resolve_field
def _ts(y, m, d):
return datetime(y, m, d, tzinfo=UTC)
def test_resolve_field_prefers_trust_then_recency():
cands = [
Candidate(value="A", source="off", trust=0.7, fetched_at=_ts(2024, 1, 1)),
Candidate(value="B", source="gs1", trust=0.9, fetched_at=_ts(2023, 1, 1)),
]
res = resolve_field(cands)
assert res is not None
assert res.value == "B"
assert res.source == "gs1"
def test_resolve_field_recency_tiebreak_on_equal_trust():
cands = [
Candidate(value="old", source="a", trust=0.7, fetched_at=_ts(2023, 1, 1)),
Candidate(value="new", source="b", trust=0.7, fetched_at=_ts(2024, 6, 1)),
]
assert resolve_field(cands).value == "new"
def test_resolve_field_skips_empty():
cands = [
Candidate(value="", source="a", trust=0.99),
Candidate(value=None, source="b", trust=0.99),
Candidate(value="kept", source="c", trust=0.1),
]
assert resolve_field(cands).value == "kept"
assert resolve_field([Candidate(value="", source="a")]) is None
def test_merge_records_provenance():
records = [
{
"name": Candidate("Water", "off", 0.7, _ts(2024, 1, 1)),
"brand": Candidate("", "off", 0.7),
},
{
"brand": Candidate("Acme", "gs1", 0.9, _ts(2024, 2, 1)),
"gtin": Candidate("123", "gs1", 0.9),
},
]
merged = merge_records(records)
assert merged.values["name"] == "Water"
assert merged.values["brand"] == "Acme"
assert merged.values["gtin"] == "123"
assert merged.provenance["brand"] == "gs1"
assert merged.provenance["name"] == "off"
def test_merge_records_accepts_bare_values():
merged = merge_records([{"x": 1}, {"x": 2}])
# both bare -> trust tie, no timestamps -> first max() wins deterministically
assert merged.values["x"] in (1, 2)
assert merged.provenance["x"] == "unknown"
+44
View File
@@ -0,0 +1,44 @@
import httpx
from opengoods.adapters.openfoodfacts import OpenFoodFactsAdapter
def _product(code, lm):
return {"code": code, "product_name": f"P{code}", "last_modified_t": lm}
def _adapter(pages):
"""Build an adapter whose search endpoint serves the given pages."""
def handler(request: httpx.Request) -> httpx.Response:
page = int(request.url.params.get("page", "1"))
products = pages.get(page, [])
return httpx.Response(200, json={"products": products, "page": page})
client = httpx.Client(transport=httpx.MockTransport(handler))
return OpenFoodFactsAdapter(client=client, min_interval=0)
def test_incremental_yields_only_newer_and_stops_at_watermark():
pages = {
1: [_product("1", 300), _product("2", 250), _product("3", 100)],
}
adapter = _adapter(pages)
got = list(adapter.fetch_modified_since(200, page_size=3, max_pages=5))
codes = [p["code"] for p in got]
assert codes == ["1", "2"] # 100 <= 200 stops iteration
def test_incremental_paginates_until_short_page():
pages = {
1: [_product("1", 900), _product("2", 800)],
2: [_product("3", 700)], # short page -> stop after
}
adapter = _adapter(pages)
got = list(adapter.fetch_modified_since(0, page_size=2, max_pages=5))
assert [p["code"] for p in got] == ["1", "2", "3"]
def test_incremental_empty_first_page():
adapter = _adapter({1: []})
assert list(adapter.fetch_modified_since(0, page_size=10, max_pages=3)) == []
+38
View File
@@ -0,0 +1,38 @@
from opengoods.etl.quality import (
COMPLETENESS_FIELDS,
agreement_from_sources,
completeness,
freshness_from_age,
score,
)
def test_completeness_bounds():
assert completeness(set()) == 0.0
assert completeness(set(COMPLETENESS_FIELDS)) == 1.0
half = set(list(COMPLETENESS_FIELDS)[: len(COMPLETENESS_FIELDS) // 2])
assert 0.0 < completeness(half) < 1.0
def test_agreement_from_sources():
assert agreement_from_sources(0) == 0.5
assert agreement_from_sources(1) == 0.5
assert agreement_from_sources(2) == 0.8
assert agreement_from_sources(5) == 1.0
def test_freshness_from_age():
assert freshness_from_age(None) == 0.5
assert freshness_from_age(1) == 1.0
assert freshness_from_age(100) == 0.8
assert freshness_from_age(300) == 0.6
assert freshness_from_age(700) == 0.4
assert freshness_from_age(5000) == 0.2
def test_score_weighted_sum_and_bounds():
assert score(completeness_score=0, source_trust=0, agreement=0, freshness=0) == 0.0
assert score(completeness_score=1, source_trust=1, agreement=1, freshness=1) == 1.0
# 0.4*1 + 0.3*0.5 + 0.2*0.5 + 0.1*1 = 0.75
got = score(completeness_score=1.0, source_trust=0.5, agreement=0.5, freshness=1.0)
assert got == 0.75
+21
View File
@@ -0,0 +1,21 @@
import json
from pathlib import Path
from opengoods.etl.load import ensure_source, load_record
from opengoods.etl.quality import compute_quality
from opengoods.etl.transform import transform
FIXTURE = json.loads((Path(__file__).parent / "fixtures" / "off_product.json").read_text())
def test_quality_score_set_on_load(db_conn):
source_id = ensure_source(db_conn)
rec = transform(FIXTURE)
pid = load_record(db_conn, rec, source_id, FIXTURE)
stored = float(
db_conn.execute("SELECT quality_score FROM product WHERE id = %s", (pid,)).fetchone()[0]
)
assert 0.0 < stored <= 1.0
# The persisted value matches a fresh recomputation.
assert abs(stored - compute_quality(db_conn, pid)) < 1e-9
+16
View File
@@ -0,0 +1,16 @@
from opengoods.etl.state import get_watermark, set_watermark
def test_watermark_roundtrip_and_monotonic(db_conn):
src = "test-source"
assert get_watermark(db_conn, src) == 0
set_watermark(db_conn, src, 100, stats={"loaded": 1})
assert get_watermark(db_conn, src) == 100
# A lower watermark must not rewind progress.
set_watermark(db_conn, src, 50)
assert get_watermark(db_conn, src) == 100
set_watermark(db_conn, src, 150)
assert get_watermark(db_conn, src) == 150
+1
View File
@@ -0,0 +1 @@
DROP TABLE IF EXISTS ingest_state;
+10
View File
@@ -0,0 +1,10 @@
-- M4 ingestion management: persistent per-source incremental watermark.
-- The updater reads/writes one row per source to resume incremental imports
-- (e.g. Open Food Facts `last_modified_t`) and to record run statistics.
CREATE TABLE ingest_state (
source TEXT PRIMARY KEY,
last_modified_t BIGINT NOT NULL DEFAULT 0,
last_run_at TIMESTAMPTZ,
cursor TEXT,
stats JSONB NOT NULL DEFAULT '{}'
);