Compare commits

...

10 Commits

Author SHA1 Message Date
novaalphastrikeomegaz663 e746b9cd31 chore(tools): add bypos-collector (条码批量采集器源码备份)
CI / Python (ingestion) (pull_request) Successful in 12s
CI / Migrations (postgres) (pull_request) Successful in 23s
CI / Go (api) (pull_request) Successful in 48s
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-24 02:20:34 +00:00
lixu 50124af833 Merge pull request 'feat(search+docs): 搜索升级(trgm 模糊 + 品牌/产地过滤 + 排序)+ 开发者文档 (M5b)' (#8) from devin/1781948300-m5b-search-docs into main
CI / Python (ingestion) (push) Successful in 10s
CI / Migrations (postgres) (push) Successful in 22s
CI / Go (api) (push) Successful in 44s
2026-06-20 17:40:07 +08:00
novaalphastrikeomegaz663 69a0149bbe feat(search+docs): trigram fuzzy search, brand/country filters, developer docs
CI / Python (ingestion) (pull_request) Successful in 12s
CI / Migrations (postgres) (pull_request) Successful in 22s
CI / Go (api) (pull_request) Successful in 47s
Search:
- migration 0009: trigram GIN index on brand.name + btree on country_of_origin
- SearchProducts: typo-tolerant word_similarity matching (>=0.42) on top of
  ILIKE substring + barcode; new brand/country filters; rank by
  similarity * (0.5 + quality_score). Response gains country_of_origin,
  quality_score and per-result relevance score.
- public search UI: brand/country filter inputs; show country in results

Docs:
- serve embedded OpenAPI 3 spec at GET /api/v1/openapi.json (not rate limited)
- ApiDocs page: auth + rate-limit section, updated search params/response
- docs/api.md developer guide

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-20 09:38:27 +00:00
lixu cf255e2380 Merge pull request 'feat(api): API Key + Redis 限流 + 用量统计 (M5a)' (#7) from devin/1781943842-m5a-api-key-ratelimit into main
CI / Python (ingestion) (push) Successful in 11s
CI / Migrations (postgres) (push) Successful in 23s
CI / Go (api) (push) Successful in 48s
2026-06-20 16:27:13 +08:00
novaalphastrikeomegaz663 2820823b36 feat(api): API keys + Redis rate limiting + usage stats
CI / Python (ingestion) (pull_request) Successful in 12s
CI / Migrations (postgres) (pull_request) Successful in 24s
CI / Go (api) (pull_request) Successful in 53s
Add an optional API-key layer to the public read-only API. Keys grant
higher per-minute rate limits and attribute usage; anonymous callers are
still allowed at a lower IP-based budget.

- migration 0008_api_key: api_key table (sha256 hash only, plaintext shown once)
- apikey pkg: key generation + hashing
- ratelimit pkg: Redis fixed-window limiter + per-key usage counters; fails open
- public API middleware: X-API-Key / Bearer auth, X-RateLimit-* headers, 429+Retry-After
- admin: issue/list/revoke keys + usage view (API + UI tab)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-20 08:24:07 +00:00
lixu 7d7a8f1baf Merge pull request 'feat(ingestion): harden OFF ingestion for bulk seeding' (#6) from devin/1781938549-ingestion-resilience into main
CI / Python (ingestion) (push) Successful in 11s
CI / Migrations (postgres) (push) Successful in 22s
CI / Go (api) (push) Successful in 36s
2026-06-20 15:54:07 +08:00
lixu e34e007f28 Merge pull request 'feat(barcode): 多条码管理(迁移 + GTIN 校验 + 后台/公开 API + 后台 UI)' (#5) from devin/1781934649-multi-barcode into main
CI / Python (ingestion) (push) Successful in 11s
CI / Migrations (postgres) (push) Successful in 22s
CI / Go (api) (push) Successful in 37s
2026-06-20 15:53:43 +08:00
novaalphastrikeomegaz663 c090bcdc9b style(ingest): ruff format country-seed tests
CI / Python (ingestion) (pull_request) Successful in 9s
CI / Migrations (postgres) (pull_request) Successful in 13s
CI / Go (api) (pull_request) Successful in 29s
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-20 07:32:30 +00:00
novaalphastrikeomegaz663 2255243081 feat(ingest): country-focused seeding (collect domestic CN products)
CI / Python (ingestion) (pull_request) Failing after 6s
CI / Migrations (postgres) (pull_request) Successful in 16s
CI / Go (api) (pull_request) Failing after 11m35s
Add a market-focused seeding path so the catalogue can be built from
domestic products rather than the English-heavy global default:

- adapter.fetch_by_country(country): OFF search filtered by
  countries_tags_en, sorted by unique_scans_n (most-scanned first),
  de-duplicated across pages since OFF popularity ordering is unstable.
- is_cn_gs1(code): True for GS1-China company prefixes (690-699),
  i.e. genuinely domestic items vs. imports merely sold in China.
- seed_off --country <slug> [--domestic-only] [--page-size/--max-pages]:
  e.g. 'seed_off --country china --domestic-only' loads only 69x
  barcodes.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-20 07:14:45 +00:00
novaalphastrikeomegaz663 044c870df7 feat(ingestion): harden OFF ingestion for bulk seeding
CI / Go (api) (pull_request) Failing after 22s
CI / Python (ingestion) (pull_request) Successful in 14s
CI / Migrations (postgres) (pull_request) Failing after 18s
- Add retry/backoff (429 + 5xx, Retry-After aware) to the OFF adapter so
  transient API errors no longer abort a run.
- Clamp bounded text fields (serving_size, net_content_unit,
  country_of_origin) to their column widths in transform; long OFF values
  previously raised StringDataRightTruncation and rolled back the batch.
- Load each record inside a savepoint (load_record_safe) so one malformed
  source record is skipped instead of aborting the whole import; jobs now
  report an errored count.
- Tests for retry behaviour, serving_size clamping, and per-record isolation.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-06-20 06:55:56 +00:00
48 changed files with 3152 additions and 86 deletions
+16 -3
View File
@@ -4,9 +4,10 @@ import Login from "./components/Login";
import ProductList from "./components/ProductList";
import ProductDetail from "./components/ProductDetail";
import SubmissionsPage from "./components/SubmissionsPage";
import { Inbox, LogOut, Package } from "lucide-react";
import ApiKeysPage from "./components/ApiKeysPage";
import { Inbox, KeyRound, LogOut, Package } from "lucide-react";
type Tab = "products" | "submissions";
type Tab = "products" | "submissions" | "keys";
type View = { name: "list" } | { name: "detail"; id: string };
export default function App() {
@@ -99,6 +100,16 @@ export default function App() {
</span>
)}
</button>
<button
onClick={() => setTab("keys")}
className={`px-3 py-1.5 rounded-md flex items-center gap-1.5 ${
tab === "keys"
? "bg-emerald-50 text-emerald-700"
: "text-gray-600 hover:bg-gray-100"
}`}
>
<KeyRound className="h-4 w-4" /> API
</button>
</nav>
</div>
<div className="flex items-center gap-4 text-sm text-gray-600">
@@ -112,7 +123,9 @@ export default function App() {
</div>
</header>
<main className="flex-1 overflow-auto p-6">
{tab === "submissions" ? (
{tab === "keys" ? (
<ApiKeysPage />
) : tab === "submissions" ? (
<SubmissionsPage onPending={setPending} />
) : view.name === "list" ? (
<ProductList onOpen={(id) => setView({ name: "detail", id })} />
+14
View File
@@ -123,4 +123,18 @@ export const api = {
method: "POST",
body: JSON.stringify({ note }),
}),
listApiKeys: () =>
request<{ items: import("./types").ApiKey[] }>("/keys"),
createApiKey: (body: {
name: string;
owner_email?: string;
tier?: string;
rate_limit_per_min?: number;
}) =>
request<{ key: string; item: import("./types").ApiKey; warning: string }>(
"/keys",
{ method: "POST", body: JSON.stringify(body) },
),
revokeApiKey: (id: string) =>
request<{ status: string }>(`/keys/${id}`, { method: "DELETE" }),
};
@@ -0,0 +1,278 @@
import { useEffect, useState } from "react";
import { api, ApiError } from "../api";
import type { ApiKey } from "../types";
import { Copy, KeyRound, Plus, Trash2 } from "lucide-react";
const TIERS = [
{ key: "free", label: "免费 (free)", rate: 120 },
{ key: "partner", label: "合作方 (partner)", rate: 600 },
{ key: "internal", label: "内部 (internal)", rate: 6000 },
];
function tierLabel(tier: string): string {
return TIERS.find((t) => t.key === tier)?.label ?? tier;
}
export default function ApiKeysPage() {
const [rows, setRows] = useState<ApiKey[]>([]);
const [error, setError] = useState("");
const [creating, setCreating] = useState(false);
const [newKey, setNewKey] = useState<string | null>(null);
async function load() {
setError("");
try {
const res = await api.listApiKeys();
setRows(res.items);
} catch (e) {
setError(e instanceof ApiError ? e.message : "加载失败");
}
}
useEffect(() => {
load();
}, []);
async function revoke(id: string, name: string) {
if (!confirm(`确认吊销密钥「${name}」?使用该密钥的请求将立即被拒绝。`)) return;
try {
await api.revokeApiKey(id);
await load();
} catch (e) {
setError(e instanceof ApiError ? e.message : "操作失败");
}
}
return (
<div className="max-w-4xl">
<div className="flex items-center justify-between mb-4">
<div>
<h2 className="text-lg font-semibold text-gray-800 flex items-center gap-2">
<KeyRound className="h-5 w-5 text-emerald-600" /> API
</h2>
<p className="text-sm text-gray-500 mt-1">
API
</p>
</div>
<button
onClick={() => setCreating(true)}
className="px-4 py-2 rounded-lg bg-emerald-600 text-white text-sm font-medium hover:bg-emerald-700 flex items-center gap-1.5"
>
<Plus className="h-4 w-4" />
</button>
</div>
{error && (
<div className="mb-3 bg-red-50 text-red-700 text-sm rounded px-4 py-2">{error}</div>
)}
{newKey && (
<div className="mb-4 bg-amber-50 border border-amber-200 rounded-lg p-4">
<div className="text-sm font-medium text-amber-800 mb-1">
</div>
<div className="flex items-center gap-2">
<code className="flex-1 bg-white border rounded px-3 py-2 text-sm break-all">
{newKey}
</code>
<button
onClick={() => navigator.clipboard?.writeText(newKey)}
className="px-3 py-2 rounded border text-sm text-gray-600 hover:bg-gray-50 flex items-center gap-1"
>
<Copy className="h-4 w-4" />
</button>
<button
onClick={() => setNewKey(null)}
className="px-3 py-2 rounded text-sm text-gray-500 hover:bg-gray-100"
>
</button>
</div>
</div>
)}
{creating && (
<CreateKeyForm
onClose={() => setCreating(false)}
onCreated={(plaintext) => {
setCreating(false);
setNewKey(plaintext);
load();
}}
/>
)}
<div className="bg-white border rounded-lg overflow-hidden">
<table className="w-full text-sm">
<thead className="bg-gray-50 text-gray-500 text-left">
<tr>
<th className="px-4 py-2 font-medium"></th>
<th className="px-4 py-2 font-medium"></th>
<th className="px-4 py-2 font-medium"></th>
<th className="px-4 py-2 font-medium">/</th>
<th className="px-4 py-2 font-medium">(/)</th>
<th className="px-4 py-2 font-medium"></th>
<th className="px-4 py-2 font-medium"></th>
</tr>
</thead>
<tbody className="divide-y">
{rows.length === 0 ? (
<tr>
<td colSpan={7} className="px-4 py-8 text-center text-gray-400">
</td>
</tr>
) : (
rows.map((k) => (
<tr key={k.id} className={k.revoked_at ? "opacity-50" : ""}>
<td className="px-4 py-2 text-gray-800">
{k.name}
{k.owner_email && (
<span className="block text-xs text-gray-400">{k.owner_email}</span>
)}
</td>
<td className="px-4 py-2 text-gray-500">
<code>{k.key_prefix}</code>
</td>
<td className="px-4 py-2 text-gray-600">{tierLabel(k.tier)}</td>
<td className="px-4 py-2 text-gray-600">{k.rate_limit_per_min}</td>
<td className="px-4 py-2 text-gray-600">
{k.usage.today} / {k.usage.total}
</td>
<td className="px-4 py-2">
{k.revoked_at ? (
<span className="text-xs rounded px-2 py-0.5 bg-red-50 text-red-700">
</span>
) : (
<span className="text-xs rounded px-2 py-0.5 bg-emerald-50 text-emerald-700">
</span>
)}
</td>
<td className="px-4 py-2 text-right">
{!k.revoked_at && (
<button
onClick={() => revoke(k.id, k.name)}
className="text-gray-400 hover:text-red-600"
title="吊销"
>
<Trash2 className="h-4 w-4" />
</button>
)}
</td>
</tr>
))
)}
</tbody>
</table>
</div>
</div>
);
}
function CreateKeyForm({
onClose,
onCreated,
}: {
onClose: () => void;
onCreated: (plaintext: string) => void;
}) {
const [name, setName] = useState("");
const [ownerEmail, setOwnerEmail] = useState("");
const [tier, setTier] = useState("free");
const [rate, setRate] = useState(120);
const [busy, setBusy] = useState(false);
const [error, setError] = useState("");
function pickTier(t: string) {
setTier(t);
const def = TIERS.find((x) => x.key === t);
if (def) setRate(def.rate);
}
async function submit() {
if (!name.trim()) {
setError("名称不能为空");
return;
}
setBusy(true);
setError("");
try {
const res = await api.createApiKey({
name: name.trim(),
owner_email: ownerEmail.trim() || undefined,
tier,
rate_limit_per_min: rate,
});
onCreated(res.key);
} catch (e) {
setError(e instanceof ApiError ? e.message : "创建失败");
} finally {
setBusy(false);
}
}
return (
<div className="mb-4 bg-white border rounded-lg p-5">
<h3 className="font-medium text-gray-700 mb-3"></h3>
{error && <div className="mb-3 bg-red-50 text-red-700 text-sm rounded px-3 py-2">{error}</div>}
<div className="grid grid-cols-2 gap-4">
<label className="block">
<span className="text-xs text-gray-500"> *</span>
<input
className="w-full border rounded-md px-3 py-2 text-sm mt-1"
value={name}
onChange={(e) => setName(e.target.value)}
placeholder="例如:我的 App / 合作方 X"
/>
</label>
<label className="block">
<span className="text-xs text-gray-500"></span>
<input
className="w-full border rounded-md px-3 py-2 text-sm mt-1"
value={ownerEmail}
onChange={(e) => setOwnerEmail(e.target.value)}
placeholder="owner@example.com"
/>
</label>
<label className="block">
<span className="text-xs text-gray-500"></span>
<select
className="w-full border rounded-md px-3 py-2 text-sm mt-1 bg-white"
value={tier}
onChange={(e) => pickTier(e.target.value)}
>
{TIERS.map((t) => (
<option key={t.key} value={t.key}>
{t.label}
</option>
))}
</select>
</label>
<label className="block">
<span className="text-xs text-gray-500">/</span>
<input
type="number"
min={1}
className="w-full border rounded-md px-3 py-2 text-sm mt-1"
value={rate}
onChange={(e) => setRate(Math.max(1, parseInt(e.target.value || "1", 10)))}
/>
</label>
</div>
<div className="mt-4 flex gap-2">
<button
onClick={submit}
disabled={busy}
className="px-4 py-2 rounded bg-emerald-600 text-white text-sm hover:bg-emerald-700 disabled:opacity-60"
>
</button>
<button onClick={onClose} className="px-4 py-2 rounded border text-sm text-gray-600">
</button>
</div>
</div>
);
}
+19
View File
@@ -137,6 +137,25 @@ export interface SubmissionDetail {
existing_product?: ProductDetail;
}
export interface ApiKeyUsage {
total: number;
today: number;
last_used_at?: number | null;
}
export interface ApiKey {
id: string;
name: string;
key_prefix: string;
owner_email: string | null;
tier: string;
rate_limit_per_min: number;
revoked_at: string | null;
created_by: string | null;
created_at: string;
usage: ApiKeyUsage;
}
export const FIELD_LABELS: Record<string, string> = {
name: "名称",
gtin: "条码",
+4 -1
View File
@@ -17,6 +17,7 @@ import (
"github.com/baicai2026-baicai/goods/api/internal/adminstore"
"github.com/baicai2026-baicai/goods/api/internal/adminweb"
"github.com/baicai2026-baicai/goods/api/internal/auth"
"github.com/baicai2026-baicai/goods/api/internal/ratelimit"
)
func getenv(key, fallback string) string {
@@ -69,7 +70,9 @@ func main() {
}
authn := auth.New(username, passwordHash, secret, 12*time.Hour)
h := adminhandler.New(adminstore.New(pool), authn, basePath, adminweb.Dist())
usage := ratelimit.New(getenv("OPENGOODS_REDIS_URL", "redis://localhost:6379/0"))
h := adminhandler.New(adminstore.New(pool), authn, basePath, adminweb.Dist()).
WithUsage(usage)
srv := &http.Server{
Addr: addr,
+7 -1
View File
@@ -12,6 +12,7 @@ import (
"github.com/baicai2026-baicai/goods/api/internal/config"
"github.com/baicai2026-baicai/goods/api/internal/handler"
"github.com/baicai2026-baicai/goods/api/internal/publicweb"
"github.com/baicai2026-baicai/goods/api/internal/ratelimit"
"github.com/baicai2026-baicai/goods/api/internal/store"
)
@@ -31,7 +32,12 @@ func main() {
log.Printf("warning: database not reachable at startup: %v", err)
}
h := handler.New(store.New(pool), publicweb.Dist())
limiter := ratelimit.New(cfg.RedisURL)
if !limiter.Enabled() {
log.Print("warning: Redis not configured; public API rate limiting disabled")
}
h := handler.New(store.New(pool), publicweb.Dist()).
WithRateLimit(limiter, cfg.AnonRateLimitPerMin)
srv := &http.Server{
Addr: cfg.Addr,
+4
View File
@@ -5,13 +5,17 @@ go 1.23.4
require (
github.com/go-chi/chi/v5 v5.1.0
github.com/jackc/pgx/v5 v5.7.2
github.com/redis/go-redis/v9 v9.18.0
golang.org/x/crypto v0.31.0
)
require (
github.com/cespare/xxhash/v2 v2.3.0 // indirect
github.com/dgryski/go-rendezvous v0.0.0-20200823014737-9f7001d12a5f // indirect
github.com/jackc/pgpassfile v1.0.0 // indirect
github.com/jackc/pgservicefile v0.0.0-20240606120523-5a60cdf6a761 // indirect
github.com/jackc/puddle/v2 v2.2.2 // indirect
go.uber.org/atomic v1.11.0 // indirect
golang.org/x/sync v0.10.0 // indirect
golang.org/x/text v0.21.0 // indirect
)
+16
View File
@@ -1,6 +1,14 @@
github.com/bsm/ginkgo/v2 v2.12.0 h1:Ny8MWAHyOepLGlLKYmXG4IEkioBysk6GpaRTLC8zwWs=
github.com/bsm/ginkgo/v2 v2.12.0/go.mod h1:SwYbGRRDovPVboqFv0tPTcG1sN61LM1Z4ARdbAV9g4c=
github.com/bsm/gomega v1.27.10 h1:yeMWxP2pV2fG3FgAODIY8EiRE3dy0aeFYt4l7wh6yKA=
github.com/bsm/gomega v1.27.10/go.mod h1:JyEr/xRbxbtgWNi8tIEVPUYZ5Dzef52k01W3YH0H+O0=
github.com/cespare/xxhash/v2 v2.3.0 h1:UL815xU9SqsFlibzuggzjXhog7bL6oX9BbNZnL2UFvs=
github.com/cespare/xxhash/v2 v2.3.0/go.mod h1:VGX0DQ3Q6kWi7AoAeZDth3/j3BFtOZR5XLFGgcrjCOs=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dgryski/go-rendezvous v0.0.0-20200823014737-9f7001d12a5f h1:lO4WD4F/rVNCu3HqELle0jiPLLBs70cWOduZpkS1E78=
github.com/dgryski/go-rendezvous v0.0.0-20200823014737-9f7001d12a5f/go.mod h1:cuUVRXasLTGF7a8hSLbxyZXjz+1KgoB3wDUb6vlszIc=
github.com/go-chi/chi/v5 v5.1.0 h1:acVI1TYaD+hhedDJ3r54HyA6sExp3HfXq7QWEEY/xMw=
github.com/go-chi/chi/v5 v5.1.0/go.mod h1:DslCQbL2OYiznFReuXYUmQ2hGd1aDpCnlMNITLSKoi8=
github.com/jackc/pgpassfile v1.0.0 h1:/6Hmqy13Ss2zCq62VdNG8tM1wchn8zjSGOBJ6icpsIM=
@@ -11,13 +19,21 @@ github.com/jackc/pgx/v5 v5.7.2 h1:mLoDLV6sonKlvjIEsV56SkWNCnuNv531l94GaIzO+XI=
github.com/jackc/pgx/v5 v5.7.2/go.mod h1:ncY89UGWxg82EykZUwSpUKEfccBGGYq1xjrOpsbsfGQ=
github.com/jackc/puddle/v2 v2.2.2 h1:PR8nw+E/1w0GLuRFSmiioY6UooMp6KJv0/61nB7icHo=
github.com/jackc/puddle/v2 v2.2.2/go.mod h1:vriiEXHvEE654aYKXXjOvZM39qJ0q+azkZFrfEOc3H4=
github.com/klauspost/cpuid/v2 v2.0.9 h1:lgaqFMSdTdQYdZ04uHyN2d/eKdOMyi2YLSvlQIBFYa4=
github.com/klauspost/cpuid/v2 v2.0.9/go.mod h1:FInQzS24/EEf25PyTYn52gqo7WaD8xa0213Md/qVLRg=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/redis/go-redis/v9 v9.18.0 h1:pMkxYPkEbMPwRdenAzUNyFNrDgHx9U+DrBabWNfSRQs=
github.com/redis/go-redis/v9 v9.18.0/go.mod h1:k3ufPphLU5YXwNTUcCRXGxUoF1fqxnhFQmscfkCoDA0=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI=
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.8.1 h1:w7B6lhMri9wdJUVmEZPGGhZzrYTPvgJArz7wNPgYKsk=
github.com/stretchr/testify v1.8.1/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o6fzry7u4=
github.com/zeebo/xxh3 v1.0.2 h1:xZmwmqxHZA8AI603jOQ0tMqmBr9lPeFwGg6d+xy9DC0=
github.com/zeebo/xxh3 v1.0.2/go.mod h1:5NWz9Sef7zIDm2JHfFlcQvNekmcEl9ekUZQQKCYaDcA=
go.uber.org/atomic v1.11.0 h1:ZvwS0R+56ePWxUNi+Atn9dWONBPp/AUETXlHW0DxSjE=
go.uber.org/atomic v1.11.0/go.mod h1:LUxbIzbOniOlMKjJjyPfpl4v+PKK2cNJn91OQbhoJI0=
golang.org/x/crypto v0.31.0 h1:ihbySMvVjLAeSH1IbfcRTkD/iNscyz8rGzjF/E5hV6U=
golang.org/x/crypto v0.31.0/go.mod h1:kDsLvtWBEx7MV9tJOj9bnXsPbxwJQ6csT/x4KIN4Ssk=
golang.org/x/sync v0.10.0 h1:3NQrjDixjgGwUOCaF8w2+VYHv0Ve/vGYSbdkTa98gmQ=
+70
View File
@@ -0,0 +1,70 @@
package adminhandler
import (
"encoding/json"
"net/http"
"strings"
"github.com/go-chi/chi/v5"
"github.com/baicai2026-baicai/goods/api/internal/adminstore"
"github.com/baicai2026-baicai/goods/api/internal/auth"
"github.com/baicai2026-baicai/goods/api/internal/ratelimit"
)
// apiKeyView is an issued key plus its usage counters.
type apiKeyView struct {
adminstore.APIKeyRow
Usage ratelimit.UsageStat `json:"usage"`
}
// ListAPIKeys returns all issued keys with usage stats merged in.
func (h *Handler) ListAPIKeys(w http.ResponseWriter, r *http.Request) {
keys, err := h.store.ListAPIKeys(r.Context())
if h.handleErr(w, err) {
return
}
views := make([]apiKeyView, 0, len(keys))
for _, k := range keys {
v := apiKeyView{APIKeyRow: k}
if h.usage != nil {
v.Usage = h.usage.Usage(r.Context(), k.ID)
}
views = append(views, v)
}
writeJSON(w, http.StatusOK, map[string]any{"items": views})
}
// CreateAPIKey issues a new key and returns its plaintext exactly once.
func (h *Handler) CreateAPIKey(w http.ResponseWriter, r *http.Request) {
var in adminstore.APIKeyInput
if err := json.NewDecoder(r.Body).Decode(&in); err != nil {
writeError(w, http.StatusBadRequest, "bad_request", "invalid body")
return
}
if strings.TrimSpace(in.Name) == "" {
writeError(w, http.StatusBadRequest, "bad_request", "名称不能为空")
return
}
if in.Tier != "" && in.Tier != "free" && in.Tier != "partner" && in.Tier != "internal" {
writeError(w, http.StatusBadRequest, "bad_request", "tier 取值无效")
return
}
plaintext, row, err := h.store.CreateAPIKey(r.Context(), in, auth.UserFrom(r.Context()))
if h.handleErr(w, err) {
return
}
writeJSON(w, http.StatusCreated, map[string]any{
"key": plaintext,
"item": row,
"warning": "请立即复制保存此密钥,它只显示这一次,无法再次查看。",
})
}
// RevokeAPIKey disables a key. Subsequent requests with it are rejected.
func (h *Handler) RevokeAPIKey(w http.ResponseWriter, r *http.Request) {
if err := h.store.RevokeAPIKey(r.Context(), chi.URLParam(r, "id")); h.handleErr(w, err) {
return
}
writeJSON(w, http.StatusOK, map[string]string{"status": "revoked"})
}
+13
View File
@@ -16,6 +16,7 @@ import (
"github.com/baicai2026-baicai/goods/api/internal/adminstore"
"github.com/baicai2026-baicai/goods/api/internal/auth"
"github.com/baicai2026-baicai/goods/api/internal/gtin"
"github.com/baicai2026-baicai/goods/api/internal/ratelimit"
)
// Handler holds the admin dependencies.
@@ -25,6 +26,7 @@ type Handler struct {
basePath string
spa fs.FS
submitLimit *rateLimiter
usage *ratelimit.Limiter
}
// New constructs an admin Handler. basePath is e.g. "/ping" (no trailing slash).
@@ -39,6 +41,13 @@ func New(store *adminstore.Store, authn *auth.Authenticator, basePath string, sp
}
}
// WithUsage attaches a Redis-backed limiter used to read per-key usage counters
// for the API-key management view. Optional; without it usage shows as zero.
func (h *Handler) WithUsage(l *ratelimit.Limiter) *Handler {
h.usage = l
return h
}
// Router builds the HTTP handler.
func (h *Handler) Router() http.Handler {
r := chi.NewRouter()
@@ -73,6 +82,10 @@ func (h *Handler) Router() http.Handler {
r.Get("/api/submissions/{id}", h.GetSubmission)
r.Post("/api/submissions/{id}/approve", h.ApproveSubmission)
r.Post("/api/submissions/{id}/reject", h.RejectSubmission)
r.Get("/api/keys", h.ListAPIKeys)
r.Post("/api/keys", h.CreateAPIKey)
r.Delete("/api/keys/{id}", h.RevokeAPIKey)
})
r.Handle("/*", http.HandlerFunc(h.serveSPA))
+130
View File
@@ -0,0 +1,130 @@
package adminstore
import (
"context"
"errors"
"strings"
"time"
"github.com/jackc/pgx/v5/pgconn"
"github.com/baicai2026-baicai/goods/api/internal/apikey"
)
// APIKeyRow is an admin-facing view of an issued API key (never the secret).
type APIKeyRow struct {
ID string `json:"id"`
Name string `json:"name"`
KeyPrefix string `json:"key_prefix"`
OwnerEmail *string `json:"owner_email"`
Tier string `json:"tier"`
RateLimitPerMin int `json:"rate_limit_per_min"`
RevokedAt *string `json:"revoked_at"`
CreatedBy *string `json:"created_by"`
CreatedAt string `json:"created_at"`
}
// APIKeyInput holds the fields accepted when issuing a key.
type APIKeyInput struct {
Name string `json:"name"`
OwnerEmail string `json:"owner_email"`
Tier string `json:"tier"`
RateLimitPerMin int `json:"rate_limit_per_min"`
}
// CreateAPIKey issues a new key, returning the one-time plaintext alongside the
// stored row. Only the SHA-256 hash and a short display prefix are persisted.
func (s *Store) CreateAPIKey(ctx context.Context, in APIKeyInput, createdBy string) (plaintext string, row APIKeyRow, err error) {
tier := in.Tier
if tier == "" {
tier = "free"
}
rate := in.RateLimitPerMin
if rate <= 0 {
rate = 120
}
var owner *string
if e := strings.TrimSpace(in.OwnerEmail); e != "" {
owner = &e
}
key, hash, prefix, err := apikey.Generate()
if err != nil {
return "", row, err
}
var revoked, created *time.Time
var createdByOut *string
err = s.pool.QueryRow(ctx, `
INSERT INTO api_key (name, key_prefix, key_hash, owner_email, tier, rate_limit_per_min, created_by)
VALUES ($1, $2, $3, $4, $5, $6, $7)
RETURNING id, name, key_prefix, owner_email, tier, rate_limit_per_min, revoked_at, created_by, created_at`,
strings.TrimSpace(in.Name), prefix, hash, owner, tier, rate, createdBy,
).Scan(&row.ID, &row.Name, &row.KeyPrefix, &row.OwnerEmail, &row.Tier,
&row.RateLimitPerMin, &revoked, &createdByOut, &created)
if err != nil {
return "", row, err
}
row.CreatedBy = createdByOut
if created != nil {
row.CreatedAt = created.Format(time.RFC3339)
}
return key, row, nil
}
// ListAPIKeys returns all keys (active first, newest first).
func (s *Store) ListAPIKeys(ctx context.Context) ([]APIKeyRow, error) {
rows, err := s.pool.Query(ctx, `
SELECT id, name, key_prefix, owner_email, tier, rate_limit_per_min, revoked_at, created_by, created_at
FROM api_key
ORDER BY (revoked_at IS NULL) DESC, created_at DESC`)
if err != nil {
return nil, err
}
defer rows.Close()
out := []APIKeyRow{}
for rows.Next() {
var r APIKeyRow
var revoked, created *time.Time
if err := rows.Scan(&r.ID, &r.Name, &r.KeyPrefix, &r.OwnerEmail, &r.Tier,
&r.RateLimitPerMin, &revoked, &r.CreatedBy, &created); err != nil {
return nil, err
}
if revoked != nil {
v := revoked.Format(time.RFC3339)
r.RevokedAt = &v
}
if created != nil {
r.CreatedAt = created.Format(time.RFC3339)
}
out = append(out, r)
}
return out, rows.Err()
}
// RevokeAPIKey marks a key revoked. Revoking an already-revoked or missing key
// returns ErrNotFound.
func (s *Store) RevokeAPIKey(ctx context.Context, id string) error {
tag, err := s.pool.Exec(ctx,
"UPDATE api_key SET revoked_at = now() WHERE id = $1 AND revoked_at IS NULL", id)
if err != nil {
if isInvalidUUID(err) {
return ErrNotFound
}
return err
}
if tag.RowsAffected() == 0 {
return ErrNotFound
}
return nil
}
// isInvalidUUID reports whether err is a Postgres invalid-UUID-text error,
// which happens when a non-UUID id is supplied.
func isInvalidUUID(err error) bool {
var pgErr *pgconn.PgError
if errors.As(err, &pgErr) {
return pgErr.Code == "22P02"
}
return false
}
+49
View File
@@ -0,0 +1,49 @@
// Package apikey handles generation and hashing of public-API keys.
//
// A key looks like "og_live_<random>". Only the SHA-256 hash is ever persisted;
// the plaintext is returned once at creation time and cannot be recovered.
package apikey
import (
"crypto/rand"
"crypto/sha256"
"encoding/base64"
"encoding/hex"
"strings"
)
// Prefix is the human-readable scheme prefix every key carries.
const Prefix = "og_live_"
// prefixLen is how many leading characters (including Prefix) are stored in
// api_key.key_prefix for identifying a key without revealing its secret.
const prefixLen = 12
// Generate returns a new random key (plaintext), its SHA-256 hash, and a short
// display prefix. The plaintext must be shown to the caller exactly once.
func Generate() (key, hash, displayPrefix string, err error) {
buf := make([]byte, 24)
if _, err = rand.Read(buf); err != nil {
return "", "", "", err
}
// URL-safe, no padding => stable, copy-pasteable token body.
body := base64.RawURLEncoding.EncodeToString(buf)
key = Prefix + body
hash = Hash(key)
displayPrefix = key
if len(displayPrefix) > prefixLen {
displayPrefix = displayPrefix[:prefixLen]
}
return key, hash, displayPrefix, nil
}
// Hash returns the hex-encoded SHA-256 of a key, used for storage and lookup.
func Hash(key string) string {
sum := sha256.Sum256([]byte(strings.TrimSpace(key)))
return hex.EncodeToString(sum[:])
}
// Looks like a key issued by this service (cheap pre-check before hashing).
func IsWellFormed(key string) bool {
return strings.HasPrefix(key, Prefix) && len(key) > len(Prefix)+8
}
+60
View File
@@ -0,0 +1,60 @@
package apikey
import "testing"
func TestGenerate(t *testing.T) {
key, hash, prefix, err := Generate()
if err != nil {
t.Fatalf("Generate: %v", err)
}
if !IsWellFormed(key) {
t.Fatalf("generated key not well-formed: %q", key)
}
if Hash(key) != hash {
t.Fatalf("Hash(key) != returned hash")
}
if len(prefix) != prefixLen || key[:prefixLen] != prefix {
t.Fatalf("prefix %q not a %d-char prefix of key %q", prefix, prefixLen, key)
}
if len(hash) != 64 {
t.Fatalf("hash not hex sha-256: %q", hash)
}
}
func TestGenerateUnique(t *testing.T) {
seen := map[string]bool{}
for i := 0; i < 100; i++ {
k, _, _, err := Generate()
if err != nil {
t.Fatal(err)
}
if seen[k] {
t.Fatalf("duplicate key generated: %q", k)
}
seen[k] = true
}
}
func TestHashStableAndTrimmed(t *testing.T) {
if Hash("og_live_abc") != Hash(" og_live_abc ") {
t.Fatal("Hash should ignore surrounding whitespace")
}
if Hash("a") == Hash("b") {
t.Fatal("distinct inputs must hash differently")
}
}
func TestIsWellFormed(t *testing.T) {
cases := map[string]bool{
"og_live_abcdefghijkl": true, // body longer than 8 chars
"og_live_": false, // empty body
"og_live_abc": false, // body too short
"nope_abcdefghijkl": false, // wrong prefix
"": false,
}
for in, want := range cases {
if got := IsWellFormed(in); got != want {
t.Errorf("IsWellFormed(%q) = %v, want %v", in, got, want)
}
}
}
+18 -6
View File
@@ -2,26 +2,38 @@ package config
import (
"os"
"strconv"
)
// Config holds runtime configuration for the OpenGoods API server.
// Values are read from environment variables with sensible defaults so the
// server can boot in a local Docker Compose setup without extra configuration.
type Config struct {
Addr string
DatabaseURL string
RedisURL string
Addr string
DatabaseURL string
RedisURL string
AnonRateLimitPerMin int
}
// Load reads configuration from the environment.
func Load() Config {
return Config{
Addr: getenv("OPENGOODS_ADDR", ":8080"),
DatabaseURL: getenv("OPENGOODS_DATABASE_URL", "postgres://opengoods:opengoods@localhost:5432/opengoods?sslmode=disable"),
RedisURL: getenv("OPENGOODS_REDIS_URL", "redis://localhost:6379/0"),
Addr: getenv("OPENGOODS_ADDR", ":8080"),
DatabaseURL: getenv("OPENGOODS_DATABASE_URL", "postgres://opengoods:opengoods@localhost:5432/opengoods?sslmode=disable"),
RedisURL: getenv("OPENGOODS_REDIS_URL", "redis://localhost:6379/0"),
AnonRateLimitPerMin: getenvInt("OPENGOODS_ANON_RATE_LIMIT_PER_MIN", 60),
}
}
func getenvInt(key string, fallback int) int {
if v, ok := os.LookupEnv(key); ok && v != "" {
if n, err := strconv.Atoi(v); err == nil && n > 0 {
return n
}
}
return fallback
}
func getenv(key, fallback string) string {
if v, ok := os.LookupEnv(key); ok && v != "" {
return v
+57 -16
View File
@@ -5,6 +5,7 @@
package handler
import (
_ "embed"
"encoding/json"
"errors"
"io/fs"
@@ -15,26 +16,48 @@ import (
"github.com/go-chi/chi/v5"
"github.com/go-chi/chi/v5/middleware"
"github.com/baicai2026-baicai/goods/api/internal/ratelimit"
"github.com/baicai2026-baicai/goods/api/internal/store"
)
//go:embed openapi.json
var openAPISpec []byte
// APIVersion is the current public API version prefix.
const APIVersion = "v1"
const (
defaultPageSize = 20
maxPageSize = 100
// defaultAnonLimit is the per-minute request budget for unauthenticated
// callers (identified by client IP) when none is configured.
defaultAnonLimit = 60
)
// Handler holds dependencies shared by the HTTP routes.
type Handler struct {
store *store.Store
spa fs.FS
store *store.Store
spa fs.FS
limiter *ratelimit.Limiter
anonLimit int
}
// New constructs a Handler backed by the given store. spa may be nil (JSON-only).
// Rate limiting is disabled until WithRateLimit is called.
func New(s *store.Store, spa fs.FS) *Handler {
return &Handler{store: s, spa: spa}
return &Handler{store: s, spa: spa, anonLimit: defaultAnonLimit}
}
// WithRateLimit attaches a Redis-backed limiter and the anonymous per-minute
// budget, enabling rate limiting + usage tracking on the public API routes.
// A non-positive anonPerMin keeps the default.
func (h *Handler) WithRateLimit(l *ratelimit.Limiter, anonPerMin int) *Handler {
h.limiter = l
if anonPerMin > 0 {
h.anonLimit = anonPerMin
}
return h
}
// Router builds the top-level HTTP handler with middleware and routes mounted.
@@ -47,16 +70,22 @@ func (h *Handler) Router() http.Handler {
r.Get("/healthz", h.Healthz)
r.Route("/api/"+APIVersion, func(r chi.Router) {
r.Route("/products", func(r chi.Router) {
r.Get("/barcode/{gtin}", h.ProductByBarcode)
r.Get("/search", h.SearchProducts)
r.Get("/{id}", h.ProductByID)
r.Get("/{id}/nutriments", h.ProductNutriments)
r.Get("/{id}/msrp", h.ProductMSRP)
// Machine-readable spec; not rate limited so tooling can always fetch it.
r.Get("/openapi.json", h.OpenAPI)
r.Group(func(r chi.Router) {
r.Use(h.rateLimit)
r.Route("/products", func(r chi.Router) {
r.Get("/barcode/{gtin}", h.ProductByBarcode)
r.Get("/search", h.SearchProducts)
r.Get("/{id}", h.ProductByID)
r.Get("/{id}/nutriments", h.ProductNutriments)
r.Get("/{id}/msrp", h.ProductMSRP)
})
r.Get("/brands", h.ListBrands)
r.Get("/categories", h.ListCategories)
r.Get("/sources/{id}", h.SourceByID)
})
r.Get("/brands", h.ListBrands)
r.Get("/categories", h.ListCategories)
r.Get("/sources/{id}", h.SourceByID)
})
// Public SPA (homepage + search + contribute). API routes above take
@@ -93,6 +122,12 @@ func (h *Handler) Healthz(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, map[string]string{"status": "ok"})
}
// OpenAPI serves the embedded OpenAPI 3 specification for the public API.
func (h *Handler) OpenAPI(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/json; charset=utf-8")
_, _ = w.Write(openAPISpec)
}
// ProductByBarcode returns a product by its GTIN.
func (h *Handler) ProductByBarcode(w http.ResponseWriter, r *http.Request) {
p, err := h.store.ProductByGTIN(r.Context(), chi.URLParam(r, "gtin"))
@@ -111,13 +146,19 @@ func (h *Handler) ProductByID(w http.ResponseWriter, r *http.Request) {
writeJSON(w, http.StatusOK, p)
}
// SearchProducts runs a fuzzy name search with optional category filter + paging.
// SearchProducts runs a trigram-fuzzy name search with optional
// category/brand/country filters, ranked by relevance, plus paging.
func (h *Handler) SearchProducts(w http.ResponseWriter, r *http.Request) {
q := r.URL.Query().Get("q")
category := r.URL.Query().Get("category")
qv := r.URL.Query()
filters := store.SearchFilters{
Query: strings.TrimSpace(qv.Get("q")),
Category: strings.TrimSpace(qv.Get("category")),
Brand: strings.TrimSpace(qv.Get("brand")),
Country: strings.TrimSpace(qv.Get("country")),
}
page, size := pageParams(r)
items, total, err := h.store.SearchProducts(r.Context(), q, category, size, (page-1)*size)
items, total, err := h.store.SearchProducts(r.Context(), filters, size, (page-1)*size)
if h.handleErr(w, r, err) {
return
}
+95
View File
@@ -116,6 +116,101 @@ func TestSearchProducts(t *testing.T) {
}
}
func TestSearchFuzzyAndFilters(t *testing.T) {
h, _ := newTestHandler(t)
ctx := context.Background()
dsn := os.Getenv("OPENGOODS_DATABASE_URL")
if dsn == "" {
dsn = "postgres://opengoods:opengoods@localhost:5432/opengoods?sslmode=disable"
}
pool, err := pgxpool.New(ctx, dsn)
if err != nil {
t.Skipf("no database: %v", err)
}
defer pool.Close()
_, err = pool.Exec(ctx,
"INSERT INTO brand (name, normalized_name) VALUES ('ZZ Test Brand','zz test brand') ON CONFLICT DO NOTHING")
if err != nil {
t.Fatalf("seed brand: %v", err)
}
_, err = pool.Exec(ctx, `
INSERT INTO product (name, brand_id, country_of_origin, quality_score, status)
VALUES ('ZZ Hazelnut Chocolate', (SELECT id FROM brand WHERE name='ZZ Test Brand'), 'Testland', 0.5, 'active')`)
if err != nil {
t.Fatalf("seed product: %v", err)
}
t.Cleanup(func() {
_, _ = pool.Exec(ctx, "DELETE FROM product WHERE name='ZZ Hazelnut Chocolate'")
_, _ = pool.Exec(ctx, "DELETE FROM brand WHERE name='ZZ Test Brand'")
})
decode := func(path string) []store.ProductSummary {
rec := doGET(t, h, path)
if rec.Code != http.StatusOK {
t.Fatalf("%s -> status %d", path, rec.Code)
}
var body struct {
Items []store.ProductSummary `json:"items"`
}
if err := json.NewDecoder(rec.Body).Decode(&body); err != nil {
t.Fatal(err)
}
return body.Items
}
has := func(items []store.ProductSummary, name string) *store.ProductSummary {
for i := range items {
if items[i].Name == name {
return &items[i]
}
}
return nil
}
// Typo "choclate" should fuzzy-match via word_similarity and carry a score.
got := has(decode("/api/"+APIVersion+"/products/search?q=choclate"), "ZZ Hazelnut Chocolate")
if got == nil {
t.Fatal("fuzzy query 'choclate' did not match 'ZZ Hazelnut Chocolate'")
}
if got.Score == nil || *got.Score <= 0 {
t.Fatalf("expected positive fuzzy score, got %v", got.Score)
}
// Brand filter.
if has(decode("/api/"+APIVersion+"/products/search?brand=ZZ+Test+Brand"), "ZZ Hazelnut Chocolate") == nil {
t.Fatal("brand filter did not return the product")
}
// Country filter (case-insensitive prefix).
if has(decode("/api/"+APIVersion+"/products/search?country=test"), "ZZ Hazelnut Chocolate") == nil {
t.Fatal("country filter did not return the product")
}
// Non-matching country excludes it.
if has(decode("/api/"+APIVersion+"/products/search?country=france"), "ZZ Hazelnut Chocolate") != nil {
t.Fatal("country filter 'france' should not return the product")
}
}
func TestOpenAPISpec(t *testing.T) {
h, _ := newTestHandler(t)
rec := doGET(t, h, "/api/"+APIVersion+"/openapi.json")
if rec.Code != http.StatusOK {
t.Fatalf("status = %d", rec.Code)
}
var spec struct {
OpenAPI string `json:"openapi"`
Paths map[string]any `json:"paths"`
}
if err := json.NewDecoder(rec.Body).Decode(&spec); err != nil {
t.Fatalf("openapi.json is not valid JSON: %v", err)
}
if spec.OpenAPI == "" || len(spec.Paths) == 0 {
t.Fatalf("unexpected spec: %+v", spec)
}
if _, ok := spec.Paths["/products/search"]; !ok {
t.Fatal("spec missing /products/search path")
}
}
func TestListCategories(t *testing.T) {
h, _ := newTestHandler(t)
rec := doGET(t, h, "/api/"+APIVersion+"/categories")
+90
View File
@@ -0,0 +1,90 @@
package handler
import (
"context"
"errors"
"net"
"net/http"
"strconv"
"strings"
"time"
"github.com/baicai2026-baicai/goods/api/internal/apikey"
"github.com/baicai2026-baicai/goods/api/internal/store"
)
type ctxKey int
const apiKeyIDKey ctxKey = 0
// rateLimit authenticates an optional API key and enforces a per-minute budget
// on the public API. Anonymous callers are limited by client IP at a lower
// budget; a valid key raises the budget and attributes usage. An API key that
// is present but invalid or revoked is rejected with 401. Rate-limit headers
// are set on every response; over-budget callers get 429 + Retry-After.
func (h *Handler) rateLimit(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
id := "ip:" + clientIP(r)
limit := h.anonLimit
keyID := ""
if raw := presentedKey(r); raw != "" {
if !apikey.IsWellFormed(raw) {
writeError(w, r, http.StatusUnauthorized, "invalid_api_key", "API key 格式无效")
return
}
k, err := h.store.APIKeyByHash(r.Context(), apikey.Hash(raw))
if errors.Is(err, store.ErrNotFound) {
writeError(w, r, http.StatusUnauthorized, "invalid_api_key", "API key 无效或已吊销")
return
}
if err != nil {
writeError(w, r, http.StatusInternalServerError, "internal_error", "internal server error")
return
}
keyID = k.ID
limit = k.RateLimitPerMin
id = "key:" + k.ID
}
res := h.limiter.Allow(r.Context(), id, limit, time.Minute)
w.Header().Set("X-RateLimit-Limit", strconv.Itoa(res.Limit))
w.Header().Set("X-RateLimit-Remaining", strconv.Itoa(res.Remaining))
w.Header().Set("X-RateLimit-Reset", strconv.FormatInt(res.ResetUnix, 10))
if !res.Allowed {
retry := res.ResetUnix - time.Now().Unix()
if retry < 1 {
retry = 1
}
w.Header().Set("Retry-After", strconv.FormatInt(retry, 10))
writeError(w, r, http.StatusTooManyRequests, "rate_limited", "请求过于频繁,请稍后再试")
return
}
if keyID != "" {
h.limiter.RecordUsage(r.Context(), keyID)
next.ServeHTTP(w, r.WithContext(context.WithValue(r.Context(), apiKeyIDKey, keyID)))
return
}
next.ServeHTTP(w, r)
})
}
// presentedKey extracts an API key from the X-API-Key header or a Bearer token.
func presentedKey(r *http.Request) string {
if v := strings.TrimSpace(r.Header.Get("X-API-Key")); v != "" {
return v
}
if v := r.Header.Get("Authorization"); strings.HasPrefix(v, "Bearer ") {
return strings.TrimSpace(strings.TrimPrefix(v, "Bearer "))
}
return ""
}
// clientIP returns the caller IP, preferring chi's RealIP-normalized RemoteAddr.
func clientIP(r *http.Request) string {
if host, _, err := net.SplitHostPort(r.RemoteAddr); err == nil {
return host
}
return r.RemoteAddr
}
+120
View File
@@ -0,0 +1,120 @@
package handler
import (
"context"
"fmt"
"net/http"
"net/http/httptest"
"os"
"testing"
"time"
"github.com/jackc/pgx/v5/pgxpool"
"github.com/baicai2026-baicai/goods/api/internal/apikey"
"github.com/baicai2026-baicai/goods/api/internal/ratelimit"
"github.com/baicai2026-baicai/goods/api/internal/store"
)
// newRateLimitedHandler builds a handler backed by the test DB and a live Redis
// limiter, plus a freshly issued API key with the given per-minute limit. It
// skips when either backend is unavailable.
func newRateLimitedHandler(t *testing.T, keyLimit int) (h *Handler, plaintextKey string) {
t.Helper()
dsn := os.Getenv("OPENGOODS_DATABASE_URL")
if dsn == "" {
dsn = "postgres://opengoods:opengoods@localhost:5432/opengoods?sslmode=disable"
}
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Second)
defer cancel()
pool, err := pgxpool.New(ctx, dsn)
if err != nil {
t.Skipf("no database: %v", err)
}
if err := pool.Ping(ctx); err != nil {
pool.Close()
t.Skipf("database not reachable: %v", err)
}
var hasTable bool
if err := pool.QueryRow(ctx, "SELECT to_regclass('public.api_key') IS NOT NULL").Scan(&hasTable); err != nil || !hasTable {
pool.Close()
t.Skip("migrations not applied (api_key missing)")
}
redisURL := os.Getenv("OPENGOODS_REDIS_URL")
if redisURL == "" {
redisURL = "redis://localhost:6379/0"
}
limiter := ratelimit.New(redisURL)
pingCtx, pingCancel := context.WithTimeout(context.Background(), time.Second)
defer pingCancel()
if err := limiter.Ping(pingCtx); err != nil {
pool.Close()
t.Skipf("redis not reachable: %v", err)
}
key, hash, prefix, err := apikey.Generate()
if err != nil {
pool.Close()
t.Fatal(err)
}
name := fmt.Sprintf("test-key-%d", time.Now().UnixNano())
if _, err := pool.Exec(context.Background(),
`INSERT INTO api_key (name, key_prefix, key_hash, rate_limit_per_min) VALUES ($1,$2,$3,$4)`,
name, prefix, hash, keyLimit); err != nil {
pool.Close()
t.Fatalf("insert api_key: %v", err)
}
t.Cleanup(func() {
_, _ = pool.Exec(context.Background(), "DELETE FROM api_key WHERE key_hash=$1", hash)
pool.Close()
})
return New(store.New(pool), nil).WithRateLimit(limiter, 60), key
}
func TestRateLimitHeadersAndKeyAuth(t *testing.T) {
h, key := newRateLimitedHandler(t, 100)
req := httptest.NewRequest(http.MethodGet, "/api/"+APIVersion+"/categories", nil)
req.Header.Set("X-API-Key", key)
rec := httptest.NewRecorder()
h.Router().ServeHTTP(rec, req)
if rec.Code != http.StatusOK {
t.Fatalf("status = %d, body = %s", rec.Code, rec.Body.String())
}
if got := rec.Header().Get("X-RateLimit-Limit"); got != "100" {
t.Fatalf("X-RateLimit-Limit = %q, want 100 (key limit)", got)
}
if rec.Header().Get("X-RateLimit-Remaining") == "" {
t.Fatal("missing X-RateLimit-Remaining header")
}
}
func TestInvalidKeyRejected(t *testing.T) {
h, _ := newRateLimitedHandler(t, 100)
req := httptest.NewRequest(http.MethodGet, "/api/"+APIVersion+"/categories", nil)
req.Header.Set("X-API-Key", "og_live_thiskeydoesnotexist123456")
rec := httptest.NewRecorder()
h.Router().ServeHTTP(rec, req)
if rec.Code != http.StatusUnauthorized {
t.Fatalf("status = %d, want 401; body = %s", rec.Code, rec.Body.String())
}
}
func TestRateLimitExceeded(t *testing.T) {
h, key := newRateLimitedHandler(t, 1)
do := func() int {
req := httptest.NewRequest(http.MethodGet, "/api/"+APIVersion+"/categories", nil)
req.Header.Set("X-API-Key", key)
rec := httptest.NewRecorder()
h.Router().ServeHTTP(rec, req)
return rec.Code
}
if code := do(); code != http.StatusOK {
t.Fatalf("first request status = %d, want 200", code)
}
if code := do(); code != http.StatusTooManyRequests {
t.Fatalf("second request status = %d, want 429", code)
}
}
+186
View File
@@ -0,0 +1,186 @@
{
"openapi": "3.0.3",
"info": {
"title": "OpenGoods / 天工商品档案公共仓 API",
"version": "1.0.0",
"description": "Public, read-only product-facts REST API. Anonymous access is allowed at a lower per-minute rate; an optional API key grants a higher rate limit and attributes usage. No purchase or commerce endpoints by design.",
"license": { "name": "Data under each source's license (e.g. ODbL)" }
},
"servers": [{ "url": "https://goods.tangshasha.com/api/v1" }],
"tags": [
{ "name": "products" },
{ "name": "catalog" },
{ "name": "meta" }
],
"security": [{ "ApiKeyHeader": [] }, { "BearerKey": [] }, {}],
"paths": {
"/products/search": {
"get": {
"tags": ["products"],
"summary": "Search products",
"description": "Trigram-fuzzy name search (typo-tolerant) with optional category/brand/country filters, ranked by name similarity blended with data quality_score.",
"parameters": [
{ "name": "q", "in": "query", "schema": { "type": "string" }, "description": "Keyword (name or barcode); fuzzy-matched. Empty returns all, ordered by quality_score." },
{ "name": "category", "in": "query", "schema": { "type": "string" }, "description": "Category code (matches the subtree), e.g. food.beverages." },
{ "name": "brand", "in": "query", "schema": { "type": "string" }, "description": "Brand name (fuzzy)." },
{ "name": "country", "in": "query", "schema": { "type": "string" }, "description": "Country of origin (case-insensitive prefix)." },
{ "name": "page", "in": "query", "schema": { "type": "integer", "default": 1, "minimum": 1 } },
{ "name": "size", "in": "query", "schema": { "type": "integer", "default": 20, "maximum": 100 } }
],
"responses": {
"200": {
"description": "Paged search results.",
"headers": {
"X-RateLimit-Limit": { "schema": { "type": "integer" }, "description": "Max requests in the current window." },
"X-RateLimit-Remaining": { "schema": { "type": "integer" }, "description": "Remaining requests in the window." },
"X-RateLimit-Reset": { "schema": { "type": "integer" }, "description": "Unix timestamp when the window resets." }
},
"content": {
"application/json": {
"schema": {
"type": "object",
"properties": {
"items": { "type": "array", "items": { "$ref": "#/components/schemas/ProductSummary" } },
"page": { "type": "integer" },
"size": { "type": "integer" },
"total": { "type": "integer" }
}
}
}
}
},
"401": { "$ref": "#/components/responses/InvalidApiKey" },
"429": { "$ref": "#/components/responses/RateLimited" }
}
}
},
"/products/barcode/{gtin}": {
"get": {
"tags": ["products"],
"summary": "Get product by barcode (GTIN)",
"parameters": [{ "name": "gtin", "in": "path", "required": true, "schema": { "type": "string" } }],
"responses": {
"200": { "description": "Product", "content": { "application/json": { "schema": { "$ref": "#/components/schemas/Product" } } } },
"404": { "$ref": "#/components/responses/NotFound" },
"429": { "$ref": "#/components/responses/RateLimited" }
}
}
},
"/products/{id}": {
"get": {
"tags": ["products"],
"summary": "Get product detail by UUID",
"parameters": [{ "name": "id", "in": "path", "required": true, "schema": { "type": "string", "format": "uuid" } }],
"responses": {
"200": { "description": "Product", "content": { "application/json": { "schema": { "$ref": "#/components/schemas/Product" } } } },
"404": { "$ref": "#/components/responses/NotFound" }
}
}
},
"/products/{id}/nutriments": {
"get": {
"tags": ["products"],
"summary": "Get product nutriments",
"parameters": [{ "name": "id", "in": "path", "required": true, "schema": { "type": "string", "format": "uuid" } }],
"responses": { "200": { "description": "Nutriments" }, "404": { "$ref": "#/components/responses/NotFound" } }
}
},
"/products/{id}/msrp": {
"get": {
"tags": ["products"],
"summary": "Get manufacturer suggested retail price snapshots (reference only)",
"parameters": [{ "name": "id", "in": "path", "required": true, "schema": { "type": "string", "format": "uuid" } }],
"responses": { "200": { "description": "MSRP snapshots" } }
}
},
"/brands": {
"get": {
"tags": ["catalog"],
"summary": "List brands",
"parameters": [
{ "name": "page", "in": "query", "schema": { "type": "integer", "default": 1 } },
{ "name": "size", "in": "query", "schema": { "type": "integer", "default": 20, "maximum": 100 } }
],
"responses": { "200": { "description": "Paged brands" } }
}
},
"/categories": {
"get": { "tags": ["catalog"], "summary": "List the category tree", "responses": { "200": { "description": "Category tree" } } }
},
"/sources/{id}": {
"get": {
"tags": ["meta"],
"summary": "Get a data source",
"parameters": [{ "name": "id", "in": "path", "required": true, "schema": { "type": "string", "format": "uuid" } }],
"responses": { "200": { "description": "Source" }, "404": { "$ref": "#/components/responses/NotFound" } }
}
}
},
"components": {
"securitySchemes": {
"ApiKeyHeader": { "type": "apiKey", "in": "header", "name": "X-API-Key", "description": "API key, e.g. og_live_xxx. Optional." },
"BearerKey": { "type": "http", "scheme": "bearer", "description": "Authorization: Bearer og_live_xxx. Optional." }
},
"responses": {
"NotFound": {
"description": "Resource not found.",
"content": { "application/json": { "schema": { "$ref": "#/components/schemas/Error" } } }
},
"RateLimited": {
"description": "Rate limit exceeded.",
"headers": { "Retry-After": { "schema": { "type": "integer" }, "description": "Seconds to wait before retrying." } },
"content": { "application/json": { "schema": { "$ref": "#/components/schemas/Error" } } }
},
"InvalidApiKey": {
"description": "API key invalid or revoked.",
"content": { "application/json": { "schema": { "$ref": "#/components/schemas/Error" } } }
}
},
"schemas": {
"Error": {
"type": "object",
"properties": {
"error": {
"type": "object",
"properties": {
"code": { "type": "string" },
"message": { "type": "string" },
"request_id": { "type": "string" }
}
}
}
},
"ProductSummary": {
"type": "object",
"properties": {
"id": { "type": "string", "format": "uuid" },
"gtin": { "type": "string", "nullable": true },
"name": { "type": "string" },
"brand": { "type": "string", "nullable": true },
"category_path": { "type": "string", "nullable": true },
"country_of_origin": { "type": "string", "nullable": true },
"quality_score": { "type": "number", "format": "float" },
"score": { "type": "number", "format": "float", "nullable": true, "description": "Relevance (name word-similarity) when q is provided; null otherwise." }
}
},
"Product": {
"type": "object",
"properties": {
"id": { "type": "string", "format": "uuid" },
"gtin": { "type": "string", "nullable": true },
"name": { "type": "string" },
"brand": { "type": "string", "nullable": true },
"category_path": { "type": "string", "nullable": true },
"net_content_value": { "type": "number", "nullable": true },
"net_content_unit": { "type": "string", "nullable": true },
"country_of_origin": { "type": "string", "nullable": true },
"quality_score": { "type": "number", "format": "float" },
"nutriments": { "type": "object", "additionalProperties": true, "nullable": true },
"nutrition_basis": { "type": "string", "nullable": true },
"nutri_score": { "type": "string", "nullable": true },
"ingredients_text": { "type": "string", "nullable": true }
}
}
}
}
}
+132
View File
@@ -0,0 +1,132 @@
// Package ratelimit provides a Redis-backed fixed-window rate limiter and
// lightweight per-key usage counters for the public API.
//
// All state lives in Redis so it is shared across API replicas and visible to
// the admin console, and so the public server keeps its read-only contract
// against PostgreSQL. Every operation fails open: if Redis is unavailable the
// limiter allows the request rather than taking the API down.
package ratelimit
import (
"context"
"fmt"
"log"
"time"
"github.com/redis/go-redis/v9"
)
// Limiter throttles callers and records usage. A nil-backed Limiter (when Redis
// could not be configured) disables limiting and usage tracking.
type Limiter struct {
rdb *redis.Client
}
// Result describes the outcome of an Allow check and the headers to surface.
type Result struct {
Allowed bool
Limit int
Remaining int
ResetUnix int64
}
// UsageStat is the aggregated usage for a single API key.
type UsageStat struct {
Total int64 `json:"total"`
Today int64 `json:"today"`
LastUsedAt *int64 `json:"last_used_at,omitempty"`
}
// New builds a Limiter from a redis:// URL. On a parse error it logs and returns
// a fail-open limiter (Redis disabled) so the server still boots.
func New(redisURL string) *Limiter {
opt, err := redis.ParseURL(redisURL)
if err != nil {
log.Printf("ratelimit: invalid redis url %q: %v (rate limiting disabled)", redisURL, err)
return &Limiter{}
}
return &Limiter{rdb: redis.NewClient(opt)}
}
// Enabled reports whether a Redis backend is configured.
func (l *Limiter) Enabled() bool { return l != nil && l.rdb != nil }
// Ping verifies the Redis backend is reachable. Returns an error if disabled or
// unreachable.
func (l *Limiter) Ping(ctx context.Context) error {
if !l.Enabled() {
return redis.ErrClosed
}
return l.rdb.Ping(ctx).Err()
}
// Allow records a hit for id within a fixed window and reports whether the
// caller is under limit. Fails open (Allowed=true) on any Redis error.
func (l *Limiter) Allow(ctx context.Context, id string, limit int, window time.Duration) Result {
reset := func() int64 {
win := int64(window / time.Second)
if win < 1 {
win = 1
}
return (time.Now().Unix()/win + 1) * win
}
if !l.Enabled() {
return Result{Allowed: true, Limit: limit, Remaining: limit, ResetUnix: reset()}
}
win := int64(window / time.Second)
if win < 1 {
win = 1
}
bucket := time.Now().Unix() / win
key := fmt.Sprintf("rl:%s:%d", id, bucket)
n, err := l.rdb.Incr(ctx, key).Result()
if err != nil {
return Result{Allowed: true, Limit: limit, Remaining: limit, ResetUnix: (bucket + 1) * win}
}
if n == 1 {
l.rdb.Expire(ctx, key, time.Duration(win)*time.Second)
}
remaining := limit - int(n)
if remaining < 0 {
remaining = 0
}
return Result{
Allowed: int(n) <= limit,
Limit: limit,
Remaining: remaining,
ResetUnix: (bucket + 1) * win,
}
}
// RecordUsage increments total/daily counters and stamps last-used for a key.
// Best-effort: errors are ignored.
func (l *Limiter) RecordUsage(ctx context.Context, keyID string) {
if !l.Enabled() || keyID == "" {
return
}
now := time.Now()
day := now.Format("20060102")
pipe := l.rdb.Pipeline()
pipe.Incr(ctx, "usage:total:"+keyID)
dayKey := "usage:day:" + keyID + ":" + day
pipe.Incr(ctx, dayKey)
pipe.Expire(ctx, dayKey, 90*24*time.Hour)
pipe.Set(ctx, "usage:last:"+keyID, now.Unix(), 0)
_, _ = pipe.Exec(ctx)
}
// Usage reads aggregated usage for a key. Returns a zero-value stat on error.
func (l *Limiter) Usage(ctx context.Context, keyID string) UsageStat {
var st UsageStat
if !l.Enabled() || keyID == "" {
return st
}
day := time.Now().Format("20060102")
st.Total, _ = l.rdb.Get(ctx, "usage:total:"+keyID).Int64()
st.Today, _ = l.rdb.Get(ctx, "usage:day:"+keyID+":"+day).Int64()
if v, err := l.rdb.Get(ctx, "usage:last:"+keyID).Int64(); err == nil {
st.LastUsedAt = &v
}
return st
}
+94
View File
@@ -0,0 +1,94 @@
package ratelimit
import (
"context"
"fmt"
"os"
"testing"
"time"
)
// TestDisabledFailsOpen verifies that a Limiter without a Redis backend allows
// all requests and reports usage as zero rather than erroring.
func TestDisabledFailsOpen(t *testing.T) {
l := New("not-a-valid-url") // parse error => disabled
if l.Enabled() {
t.Fatal("expected limiter to be disabled for invalid url")
}
res := l.Allow(context.Background(), "x", 1, time.Minute)
if !res.Allowed || res.Remaining != 1 {
t.Fatalf("disabled limiter must fail open: %+v", res)
}
// Must not panic and must return zero usage.
l.RecordUsage(context.Background(), "k1")
if u := l.Usage(context.Background(), "k1"); u.Total != 0 {
t.Fatalf("disabled usage should be zero, got %+v", u)
}
}
// TestNilReceiverSafe ensures a nil *Limiter is safe to use (handler default).
func TestNilReceiverSafe(t *testing.T) {
var l *Limiter
if l.Enabled() {
t.Fatal("nil limiter must report disabled")
}
res := l.Allow(context.Background(), "x", 5, time.Minute)
if !res.Allowed {
t.Fatal("nil limiter must fail open")
}
l.RecordUsage(context.Background(), "k")
_ = l.Usage(context.Background(), "k")
}
func testLimiter(t *testing.T) *Limiter {
t.Helper()
url := os.Getenv("OPENGOODS_REDIS_URL")
if url == "" {
url = "redis://localhost:6379/0"
}
l := New(url)
if !l.Enabled() {
t.Skip("redis not configured")
}
ctx, cancel := context.WithTimeout(context.Background(), time.Second)
defer cancel()
if err := l.rdb.Ping(ctx).Err(); err != nil {
t.Skipf("redis not reachable: %v", err)
}
return l
}
func TestAllowFixedWindow(t *testing.T) {
l := testLimiter(t)
ctx := context.Background()
id := fmt.Sprintf("test:%d", time.Now().UnixNano())
for i := 1; i <= 2; i++ {
if res := l.Allow(ctx, id, 2, time.Minute); !res.Allowed {
t.Fatalf("request %d should be allowed: %+v", i, res)
}
}
res := l.Allow(ctx, id, 2, time.Minute)
if res.Allowed {
t.Fatalf("3rd request over limit 2 should be denied: %+v", res)
}
if res.Remaining != 0 {
t.Fatalf("remaining should be 0 when over limit, got %d", res.Remaining)
}
}
func TestRecordAndReadUsage(t *testing.T) {
l := testLimiter(t)
ctx := context.Background()
key := fmt.Sprintf("usagekey:%d", time.Now().UnixNano())
l.RecordUsage(ctx, key)
l.RecordUsage(ctx, key)
u := l.Usage(ctx, key)
if u.Total != 2 || u.Today != 2 {
t.Fatalf("expected total=2 today=2, got %+v", u)
}
if u.LastUsedAt == nil {
t.Fatal("expected last-used timestamp to be set")
}
}
+91 -22
View File
@@ -82,11 +82,22 @@ func (s *Store) ProductBarcodes(ctx context.Context, productID string) ([]Barcod
// ProductSummary is a lightweight row used in search/listing responses.
type ProductSummary struct {
ID string `json:"id"`
GTIN *string `json:"gtin"`
Name string `json:"name"`
Brand *string `json:"brand"`
CategoryPath *string `json:"category_path"`
ID string `json:"id"`
GTIN *string `json:"gtin"`
Name string `json:"name"`
Brand *string `json:"brand"`
CategoryPath *string `json:"category_path"`
Country *string `json:"country_of_origin"`
QualityScore float64 `json:"quality_score"`
Score *float64 `json:"score,omitempty"`
}
// SearchFilters bundles the optional filters accepted by SearchProducts.
type SearchFilters struct {
Query string // fuzzy name / barcode query
Category string // ltree path; matches the subtree
Brand string // fuzzy brand name
Country string // country_of_origin prefix (case-insensitive)
}
const productSelect = `
@@ -148,34 +159,67 @@ func (s *Store) ProductByID(ctx context.Context, id string) (*Product, error) {
return p, nil
}
// SearchProducts performs a fuzzy name search with optional category subtree filter.
func (s *Store) SearchProducts(ctx context.Context, q, category string, limit, offset int) ([]ProductSummary, int, error) {
// fuzzyThreshold is the minimum word_similarity for a name to be considered a
// fuzzy match. ~0.42 tolerates common typos (e.g. "choclate"→"Chocolate")
// without returning unrelated products.
const fuzzyThreshold = "0.42"
// SearchProducts runs a trigram-fuzzy name search with optional category /
// brand / country filters. When a query is present, matching is inclusive
// (substring OR trigram-similar OR barcode), and results are ranked by name
// similarity blended with quality_score so the best, most-complete records
// surface first. Without a query, results are ordered by quality_score.
func (s *Store) SearchProducts(ctx context.Context, f SearchFilters, limit, offset int) ([]ProductSummary, int, error) {
args := []any{}
where := "WHERE p.status = 'active'"
if q != "" {
args = append(args, q)
where += ` AND (p.name ILIKE '%' || $1 || '%'
qIdx := 0
if f.Query != "" {
args = append(args, f.Query)
qIdx = len(args)
q := "$" + strconv.Itoa(qIdx)
where += ` AND (p.name ILIKE '%' || ` + q + ` || '%'
OR word_similarity(` + q + `, p.name) >= ` + fuzzyThreshold + `
OR EXISTS (SELECT 1 FROM product_barcode pb
WHERE pb.product_id = p.id AND pb.gtin ILIKE '%' || $1 || '%'))`
WHERE pb.product_id = p.id AND pb.gtin ILIKE '%' || ` + q + ` || '%'))`
}
if category != "" {
args = append(args, category)
if f.Category != "" {
args = append(args, f.Category)
where += " AND c.path <@ $" + strconv.Itoa(len(args)) + "::ltree"
}
if f.Brand != "" {
args = append(args, f.Brand)
where += " AND b.name ILIKE '%' || $" + strconv.Itoa(len(args)) + " || '%'"
}
if f.Country != "" {
args = append(args, f.Country)
where += " AND p.country_of_origin ILIKE $" + strconv.Itoa(len(args)) + " || '%'"
}
from := `FROM product p
LEFT JOIN brand b ON b.id = p.brand_id
LEFT JOIN category c ON c.id = p.category_id `
countSQL := "SELECT count(*) FROM product p LEFT JOIN category c ON c.id = p.category_id " + where
var total int
if err := s.pool.QueryRow(ctx, countSQL, args...).Scan(&total); err != nil {
if err := s.pool.QueryRow(ctx, "SELECT count(*) "+from+where, args...).Scan(&total); err != nil {
return nil, 0, err
}
// Ranking: when querying, similarity drives order, multiplied by a
// quality factor floored at 0.5 so low-quality records aren't zeroed out.
scoreExpr := "NULL::real"
orderBy := "p.quality_score DESC, p.name"
if f.Query != "" {
q := "$" + strconv.Itoa(qIdx)
scoreExpr = "word_similarity(" + q + ", p.name)"
orderBy = scoreExpr + " * (0.5 + p.quality_score) DESC, p.quality_score DESC, p.name"
}
args = append(args, limit, offset)
listSQL := `
SELECT p.id, p.gtin, p.name, b.name, c.path::text
FROM product p
LEFT JOIN brand b ON b.id = p.brand_id
LEFT JOIN category c ON c.id = p.category_id ` + where +
" ORDER BY p.name LIMIT $" + strconv.Itoa(len(args)-1) + " OFFSET $" + strconv.Itoa(len(args))
listSQL := "SELECT p.id, p.gtin, p.name, b.name, c.path::text, p.country_of_origin, p.quality_score, " +
scoreExpr + " AS score " + from + where +
" ORDER BY " + orderBy +
" LIMIT $" + strconv.Itoa(len(args)-1) + " OFFSET $" + strconv.Itoa(len(args))
rows, err := s.pool.Query(ctx, listSQL, args...)
if err != nil {
@@ -186,7 +230,8 @@ LEFT JOIN category c ON c.id = p.category_id ` + where +
out := []ProductSummary{}
for rows.Next() {
var ps ProductSummary
if err := rows.Scan(&ps.ID, &ps.GTIN, &ps.Name, &ps.Brand, &ps.CategoryPath); err != nil {
if err := rows.Scan(&ps.ID, &ps.GTIN, &ps.Name, &ps.Brand, &ps.CategoryPath,
&ps.Country, &ps.QualityScore, &ps.Score); err != nil {
return nil, 0, err
}
out = append(out, ps)
@@ -305,6 +350,30 @@ func (s *Store) ListCategories(ctx context.Context) ([]Category, error) {
return out, rows.Err()
}
// APIKey is the minimal metadata the public API needs to authorize a caller.
type APIKey struct {
ID string
Name string
RateLimitPerMin int
}
// APIKeyByHash returns the active (non-revoked) key matching a SHA-256 hash,
// or ErrNotFound if no such active key exists.
func (s *Store) APIKeyByHash(ctx context.Context, hash string) (*APIKey, error) {
var k APIKey
err := s.pool.QueryRow(ctx,
`SELECT id, name, rate_limit_per_min
FROM api_key WHERE key_hash = $1 AND revoked_at IS NULL`, hash,
).Scan(&k.ID, &k.Name, &k.RateLimitPerMin)
if errors.Is(err, pgx.ErrNoRows) {
return nil, ErrNotFound
}
if err != nil {
return nil, err
}
return &k, nil
}
// Source describes a data source with its license and trust weight.
type Source struct {
ID string `json:"id"`
+117
View File
@@ -0,0 +1,117 @@
# OpenGoods 公共 API 开发者文档
天工商品档案公共仓(OpenGoods)提供**公开、只读**的商品事实 REST API:按条码/名称查询商品的客观资料(品牌、品类、净含量、产地、配料、营养成分、Nutri-Score、厂商建议零售价快照等)。返回均为 JSON(UTF-8)。**本服务不含任何购买/交易接口。**
- 基础地址:`https://goods.tangshasha.com/api/v1`
- 机器可读规范(OpenAPI 3):`GET /api/v1/openapi.json`
- 交互式文档:站点「API 调用说明」页
## 鉴权
API 默认**匿名可用**,无需任何凭证即可调用。匿名请求按来源 IP 计入一个较低的默认每分钟额度。
如需更高额度并让用量归属到你,可向运营方申请一枚 **API Key**(形如 `og_live_xxxxxxxx`),请求时二选一携带:
```bash
curl -H "X-API-Key: og_live_xxxxxxxx" \
"https://goods.tangshasha.com/api/v1/products/search?q=牛奶"
# 或
curl -H "Authorization: Bearer og_live_xxxxxxxx" \
"https://goods.tangshasha.com/api/v1/products/search?q=牛奶"
```
> 仅在创建时返回一次明文 Key,请妥善保存。服务端只存储其 SHA-256 哈希。
## 限流
采用**固定窗口**限流(每分钟)。每个响应都会回写以下响应头:
| 响应头 | 含义 |
| --- | --- |
| `X-RateLimit-Limit` | 当前窗口允许的最大请求数 |
| `X-RateLimit-Remaining` | 当前窗口剩余可用次数 |
| `X-RateLimit-Reset` | 窗口重置的 Unix 时间戳(秒) |
| `Retry-After` | 仅在超额(429)时返回,建议等待的秒数 |
- 超过额度:`429 Too Many Requests`,错误码 `rate_limited`
- Key 无效或已吊销:`401 Unauthorized`,错误码 `invalid_api_key`
## 错误格式
非 2xx 响应体统一为:
```json
{ "error": { "code": "not_found", "message": "…", "request_id": "…" } }
```
## 分页
列表类接口支持 `page`(默认 `1`)与 `size`(默认 `20`,最大 `100`),响应含 `page`/`size`/`total`
## 端点
### `GET /products/search` — 搜索商品
按名称做三元组(trigram)模糊搜索,**可容忍错别字**;支持品类/品牌/产地过滤;结果按相关度(名称相似度 × 数据质量分)排序。
| 参数 | 必填 | 说明 |
| --- | --- | --- |
| `q` | 否 | 关键词(名称/条码),模糊匹配;留空则按质量分返回全部 |
| `category` | 否 | 品类编码(含子树),如 `food.beverages` |
| `brand` | 否 | 品牌名(模糊匹配),如 `Ferrero` |
| `country` | 否 | 产地前缀(不区分大小写),如 `China` |
| `page` | 否 | 页码,默认 1 |
| `size` | 否 | 每页条数,默认 20,最大 100 |
```bash
curl "https://goods.tangshasha.com/api/v1/products/search?q=nutela&country=Italy"
```
```json
{
"items": [
{
"id": "…",
"gtin": "3017624010701",
"name": "Nutella",
"brand": "Ferrero",
"category_path": "food.snacks.chocolate",
"country_of_origin": "Italy",
"quality_score": 0.81,
"score": 0.71
}
],
"page": 1,
"size": 20,
"total": 1
}
```
`score` 为名称相关度(提供 `q` 时返回,01),未提供 `q` 时为 `null`
### `GET /products/barcode/{gtin}` — 按条码查询
```bash
curl "https://goods.tangshasha.com/api/v1/products/barcode/5449000000996"
```
### `GET /products/{id}` — 商品详情
按商品 UUID 获取完整档案(含配料、营养、添加剂、图片、MSRP 等)。
### `GET /products/{id}/nutriments` — 商品营养成分
### `GET /products/{id}/msrp` — 厂商建议零售价快照
官方建议零售价历史快照,仅供参考,不含任何购买入口。
### `GET /brands` — 品牌列表(分页)
### `GET /categories` — 品类树
### `GET /sources/{id}` — 数据来源
## 免责声明
数据可能存在误差或滞后,按「现状」提供,不构成医疗/购买建议。商品资料版权归各原始来源所有,请遵循其许可(如 OpenFoodFacts 的 ODbL),引用时请注明天工商品档案公共仓及原始来源。
+103 -6
View File
@@ -27,6 +27,9 @@ _DEFAULT_MIN_INTERVAL = 4.0
_API_URL = "https://world.openfoodfacts.org/api/v2/product/{barcode}.json"
_SEARCH_URL = "https://world.openfoodfacts.org/api/v2/search"
# HTTP statuses worth retrying: rate limiting and transient server errors.
_RETRY_STATUS = frozenset({429, 500, 502, 503, 504})
# Fields requested from the search API so a returned product can be transformed
# without an extra per-barcode round trip.
_SEARCH_FIELDS = (
@@ -46,9 +49,13 @@ class OpenFoodFactsAdapter:
self,
client: httpx.Client | None = None,
min_interval: float = _DEFAULT_MIN_INTERVAL,
max_retries: int = 4,
backoff_base: float = 2.0,
) -> None:
self._client = client or httpx.Client(headers={"User-Agent": USER_AGENT}, timeout=30.0)
self._min_interval = min_interval
self._max_retries = max_retries
self._backoff_base = backoff_base
self._last_call = 0.0
def _throttle(self) -> None:
@@ -58,11 +65,49 @@ class OpenFoodFactsAdapter:
time.sleep(wait)
self._last_call = time.monotonic()
def _get(self, url: str, params: dict | None = None) -> httpx.Response:
"""GET with throttling and retry/backoff on transient errors.
Retries on connection/timeout errors and on retryable HTTP statuses
(429 and 5xx, which OFF returns intermittently when overloaded), using
exponential backoff that honours a ``Retry-After`` header when present.
"""
last_exc: Exception | None = None
for attempt in range(self._max_retries + 1):
self._throttle()
try:
resp = self._client.get(url, params=params)
except httpx.TransportError as exc:
last_exc = exc
else:
if resp.status_code < 400 or resp.status_code not in _RETRY_STATUS:
resp.raise_for_status()
return resp
last_exc = httpx.HTTPStatusError(
f"retryable status {resp.status_code}", request=resp.request, response=resp
)
if attempt < self._max_retries:
retry_after = self._retry_after(last_exc)
time.sleep(retry_after if retry_after is not None else self._backoff_base**attempt)
assert last_exc is not None
raise last_exc
@staticmethod
def _retry_after(exc: Exception | None) -> float | None:
resp = getattr(exc, "response", None)
if resp is None:
return None
value = resp.headers.get("Retry-After")
if not value:
return None
try:
return float(value)
except ValueError:
return None
def fetch_barcode(self, barcode: str) -> dict | None:
"""Fetch a single product by barcode; return the raw `product` dict."""
self._throttle()
resp = self._client.get(_API_URL.format(barcode=barcode))
resp.raise_for_status()
resp = self._get(_API_URL.format(barcode=barcode))
payload = resp.json()
if payload.get("status") != 1:
return None
@@ -91,8 +136,7 @@ class OpenFoodFactsAdapter:
``last_modified_t`` they processed as the next watermark.
"""
for page in range(1, max_pages + 1):
self._throttle()
resp = self._client.get(
resp = self._get(
_SEARCH_URL,
params={
"fields": _SEARCH_FIELDS,
@@ -101,7 +145,6 @@ class OpenFoodFactsAdapter:
"page_size": page_size,
},
)
resp.raise_for_status()
products = resp.json().get("products") or []
if not products:
return
@@ -114,6 +157,60 @@ class OpenFoodFactsAdapter:
if reached_old or len(products) < page_size:
return
def fetch_by_country(
self,
country: str,
*,
page_size: int = 100,
max_pages: int = 10,
sort_by: str = "unique_scans_n",
) -> Iterator[dict]:
"""Yield products sold in ``country`` (an OFF ``countries_tags_en`` slug).
Used to seed a market-specific catalogue (e.g. ``china``). Results are
sorted by ``sort_by`` (default ``unique_scans_n`` so the most-scanned,
best-known products come first) and de-duplicated across pages, since
OFF's popularity ordering is not stable between page requests.
"""
seen: set[str] = set()
for page in range(1, max_pages + 1):
resp = self._get(
_SEARCH_URL,
params={
"fields": _SEARCH_FIELDS,
"countries_tags_en": country,
"sort_by": sort_by,
"page": page,
"page_size": page_size,
},
)
products = resp.json().get("products") or []
if not products:
return
new_on_page = 0
for prod in products:
code = str(prod.get("code") or "")
if code and code in seen:
continue
if code:
seen.add(code)
new_on_page += 1
yield prod
if len(products) < page_size or new_on_page == 0:
return
def is_cn_gs1(code: str | None) -> bool:
"""Return True for a GS1 China company prefix (barcodes starting 690-699).
These identify products registered with GS1 China, i.e. genuinely domestic
items, as opposed to imported goods merely tagged as sold in China.
"""
if not code:
return False
code = code.strip()
return len(code) >= 3 and code[:2] == "69" and code[2].isdigit()
def read_dump(path: str | Path) -> Iterator[dict]:
"""Yield raw product records from an OFF JSONL dump file.
+22
View File
@@ -7,6 +7,7 @@ source with field-level provenance in `product_source`.
from __future__ import annotations
import json
import logging
import os
from typing import Any
@@ -18,6 +19,8 @@ from opengoods.etl.quality import update_quality
OFF_HOMEPAGE = "https://world.openfoodfacts.org"
logger = logging.getLogger(__name__)
def default_dsn() -> str:
return os.environ.get(
@@ -196,6 +199,25 @@ def load_record(conn: psycopg.Connection, rec: dict[str, Any], source_id: str, r
return product_id
def load_record_safe(
conn: psycopg.Connection, rec: dict[str, Any], source_id: str, raw: dict
) -> bool:
"""Load one record inside a savepoint.
On success the record's writes stay in the surrounding transaction. On any
error, only this record's writes are rolled back (to the savepoint) and the
batch continues, so a single malformed source record cannot abort a large
import. Returns True if loaded, False if skipped due to an error.
"""
try:
with conn.transaction():
load_record(conn, rec, source_id, raw)
return True
except Exception as exc: # noqa: BLE001 - per-record isolation is intentional
logger.warning("skipping record gtin=%s: %s", rec.get("gtin"), exc)
return False
def _jsonable(raw: dict) -> dict:
"""Drop values that are not JSON-serializable from a raw record."""
try:
+11 -3
View File
@@ -78,6 +78,14 @@ def map_category(raw: dict) -> str | None:
return None
def _clamp(value: str | None, max_len: int) -> str | None:
"""Trim a string to fit a bounded DB column; external data length varies."""
if value is None:
return None
value = value.strip()
return value[:max_len] or None
def _clean_tags(tags: list[str] | None, prefix: str = "") -> list[str]:
out: list[str] = []
for t in tags or []:
@@ -143,16 +151,16 @@ def transform(raw: dict) -> dict | None:
"brand": brand,
"category_path": map_category(raw),
"net_content_value": net_value,
"net_content_unit": net_unit,
"net_content_unit": _clamp(net_unit, 16),
"net_content_canonical": net_canonical,
"country_of_origin": (raw.get("countries") or "").split(",")[0].strip() or None,
"country_of_origin": _clamp((raw.get("countries") or "").split(",")[0].strip() or None, 64),
"food": {
"ingredients_text": raw.get("ingredients_text") or None,
"allergens": _clean_tags(raw.get("allergens_tags")),
"additives": _clean_tags(raw.get("additives_tags")),
"nutriments": transform_nutriments(raw.get("nutriments") or {}),
"nutrition_basis": "per_100g",
"serving_size": raw.get("serving_size") or None,
"serving_size": _clamp(raw.get("serving_size") or None, 32),
"nutri_score": (raw.get("nutriscore_grade") or "").upper()[:1] or None,
},
"image_url": raw.get("image_front_url") or raw.get("image_url") or None,
+40 -9
View File
@@ -7,6 +7,11 @@ Usage:
# from a downloaded OFF JSONL dump (optionally .gz), limited to N records
python -m opengoods.jobs.seed_off --dump products.jsonl.gz --limit 1000
# market-focused: the most-scanned products sold in China, restricted to
# genuine GS1-China (69x) barcodes
python -m opengoods.jobs.seed_off --country china --domestic-only \
--max-pages 20 --limit 1000
The OFF read API is rate-limited client-side; for large imports use a dump.
"""
@@ -18,25 +23,38 @@ from collections.abc import Iterator
import psycopg
from opengoods.adapters.openfoodfacts import OpenFoodFactsAdapter, read_dump
from opengoods.etl.load import default_dsn, ensure_source, load_record
from opengoods.adapters.openfoodfacts import (
OpenFoodFactsAdapter,
is_cn_gs1,
read_dump,
)
from opengoods.etl.load import default_dsn, ensure_source, load_record_safe
from opengoods.etl.transform import transform
def _raw_records(args: argparse.Namespace) -> Iterator[dict]:
if args.dump:
records = read_dump(args.dump)
records: Iterator[dict] = read_dump(args.dump)
elif args.country:
adapter = OpenFoodFactsAdapter(min_interval=args.min_interval)
records = adapter.fetch_by_country(
args.country, page_size=args.page_size, max_pages=args.max_pages
)
else:
adapter = OpenFoodFactsAdapter(min_interval=args.min_interval)
records = adapter.fetch(args.barcodes)
for i, rec in enumerate(records):
if args.limit and i >= args.limit:
yielded = 0
for rec in records:
if args.domestic_only and not is_cn_gs1(str(rec.get("code") or "")):
continue
if args.limit and yielded >= args.limit:
break
yielded += 1
yield rec
def run(args: argparse.Namespace) -> int:
loaded = skipped = 0
loaded = skipped = errored = 0
with psycopg.connect(args.dsn, autocommit=False) as conn:
source_id = ensure_source(conn)
for raw in _raw_records(args):
@@ -44,10 +62,12 @@ def run(args: argparse.Namespace) -> int:
if rec is None:
skipped += 1
continue
load_record(conn, rec, source_id, raw)
loaded += 1
if load_record_safe(conn, rec, source_id, raw):
loaded += 1
else:
errored += 1
conn.commit()
print(f"loaded={loaded} skipped={skipped}")
print(f"loaded={loaded} skipped={skipped} errored={errored}")
return 0
@@ -56,7 +76,18 @@ def main(argv: list[str] | None = None) -> int:
src = parser.add_mutually_exclusive_group(required=True)
src.add_argument("--barcodes", nargs="+", help="barcodes to fetch via the OFF API")
src.add_argument("--dump", help="path to an OFF JSONL dump (.jsonl or .jsonl.gz)")
src.add_argument(
"--country",
help="OFF countries_tags_en slug to seed from, e.g. 'china' (most-scanned first)",
)
parser.add_argument(
"--domestic-only",
action="store_true",
help="keep only genuine GS1-China (69x) barcodes; drop imported goods",
)
parser.add_argument("--limit", type=int, default=0, help="max records to load (0 = all)")
parser.add_argument("--page-size", type=int, default=100, help="search page size")
parser.add_argument("--max-pages", type=int, default=10, help="max search pages (country mode)")
parser.add_argument("--min-interval", type=float, default=4.0, help="API throttle seconds")
parser.add_argument("--dsn", default=default_dsn(), help="PostgreSQL DSN")
return run(parser.parse_args(argv))
+11 -6
View File
@@ -17,14 +17,14 @@ import sys
import psycopg
from opengoods.adapters.openfoodfacts import SOURCE_NAME, OpenFoodFactsAdapter
from opengoods.etl.load import default_dsn, ensure_source, load_record
from opengoods.etl.load import default_dsn, ensure_source, load_record_safe
from opengoods.etl.state import get_watermark, set_watermark
from opengoods.etl.transform import transform
def run(args: argparse.Namespace) -> int:
adapter = OpenFoodFactsAdapter(min_interval=args.min_interval)
loaded = skipped = 0
loaded = skipped = errored = 0
high_watermark = 0
with psycopg.connect(args.dsn, autocommit=False) as conn:
source_id = ensure_source(conn)
@@ -38,16 +38,21 @@ def run(args: argparse.Namespace) -> int:
if rec is None:
skipped += 1
continue
load_record(conn, rec, source_id, raw)
loaded += 1
if load_record_safe(conn, rec, source_id, raw):
loaded += 1
else:
errored += 1
set_watermark(
conn,
SOURCE_NAME,
high_watermark,
stats={"loaded": loaded, "skipped": skipped, "since": since},
stats={"loaded": loaded, "skipped": skipped, "errored": errored, "since": since},
)
conn.commit()
print(f"since={since} loaded={loaded} skipped={skipped} watermark={high_watermark}")
print(
f"since={since} loaded={loaded} skipped={skipped} "
f"errored={errored} watermark={high_watermark}"
)
return 0
+23 -1
View File
@@ -11,7 +11,7 @@ from pathlib import Path
import pytest
from opengoods.etl.load import default_dsn, ensure_source, load_record
from opengoods.etl.load import default_dsn, ensure_source, load_record, load_record_safe
from opengoods.etl.transform import transform
psycopg = pytest.importorskip("psycopg")
@@ -59,3 +59,25 @@ def test_load_record_roundtrip(conn):
assert prov[0] >= 1
conn.rollback() # keep the test DB clean
def test_load_record_safe_isolates_bad_record(conn):
source_id = ensure_source(conn)
# Unique gtin so the good record is a fresh INSERT, not an upsert/update.
good = transform(FIXTURE)
good["gtin"] = "4006381333931"
assert load_record_safe(conn, good, source_id, FIXTURE) is True
after_good = conn.execute("SELECT count(*) FROM product").fetchone()[0]
# A record whose name violates NOT NULL must not abort the batch.
bad = dict(good)
bad["gtin"] = "5000112637922"
bad["name"] = None
assert load_record_safe(conn, bad, source_id, {}) is False
# The good record survived the bad one's rollback-to-savepoint.
after_bad = conn.execute("SELECT count(*) FROM product").fetchone()[0]
assert after_bad == after_good
conn.rollback() # keep the test DB clean
+53
View File
@@ -0,0 +1,53 @@
import httpx
from opengoods.adapters.openfoodfacts import OpenFoodFactsAdapter, is_cn_gs1
def _adapter(handler, **kwargs):
client = httpx.Client(transport=httpx.MockTransport(handler))
return OpenFoodFactsAdapter(client=client, min_interval=0, backoff_base=0, **kwargs)
def test_is_cn_gs1():
assert is_cn_gs1("6901234567892") # GS1 China prefix
assert is_cn_gs1("690")
assert not is_cn_gs1("3017624010701") # France
assert not is_cn_gs1("5449000000996") # Belgium
assert not is_cn_gs1("")
assert not is_cn_gs1(None)
assert not is_cn_gs1("69") # too short to carry a prefix digit
def test_fetch_by_country_passes_filter_and_paginates():
seen_params = []
def handler(request: httpx.Request) -> httpx.Response:
seen_params.append(dict(request.url.params))
page = int(request.url.params.get("page", "1"))
if page == 1:
return httpx.Response(
200,
json={"products": [{"code": "6901"}, {"code": "6902"}]},
)
return httpx.Response(200, json={"products": []})
adapter = _adapter(handler)
out = list(adapter.fetch_by_country("china", page_size=2, max_pages=5))
assert [p["code"] for p in out] == ["6901", "6902"]
assert seen_params[0]["countries_tags_en"] == "china"
assert seen_params[0]["sort_by"] == "unique_scans_n"
def test_fetch_by_country_dedupes_across_pages():
def handler(request: httpx.Request) -> httpx.Response:
page = int(request.url.params.get("page", "1"))
if page == 1:
return httpx.Response(200, json={"products": [{"code": "6901"}, {"code": "6902"}]})
if page == 2:
# OFF popularity ordering is unstable; a repeat appears on page 2
return httpx.Response(200, json={"products": [{"code": "6902"}, {"code": "6903"}]})
return httpx.Response(200, json={"products": []})
adapter = _adapter(handler)
out = [p["code"] for p in adapter.fetch_by_country("china", page_size=2, max_pages=5)]
assert out == ["6901", "6902", "6903"]
+59
View File
@@ -0,0 +1,59 @@
import httpx
import pytest
from opengoods.adapters.openfoodfacts import OpenFoodFactsAdapter
def _adapter(handler, **kwargs):
client = httpx.Client(transport=httpx.MockTransport(handler))
return OpenFoodFactsAdapter(client=client, min_interval=0, backoff_base=0, **kwargs)
def test_retries_transient_5xx_then_succeeds():
calls = {"n": 0}
def handler(request: httpx.Request) -> httpx.Response:
calls["n"] += 1
if calls["n"] < 3:
return httpx.Response(503)
return httpx.Response(200, json={"status": 1, "product": {"code": "x"}})
adapter = _adapter(handler, max_retries=4)
assert adapter.fetch_barcode("x") == {"code": "x"}
assert calls["n"] == 3 # two 503s retried, third succeeds
def test_retries_exhausted_raises():
def handler(request: httpx.Request) -> httpx.Response:
return httpx.Response(503)
adapter = _adapter(handler, max_retries=2)
with pytest.raises(httpx.HTTPStatusError):
adapter.fetch_barcode("x")
def test_non_retryable_4xx_not_retried():
calls = {"n": 0}
def handler(request: httpx.Request) -> httpx.Response:
calls["n"] += 1
return httpx.Response(404)
adapter = _adapter(handler, max_retries=4)
with pytest.raises(httpx.HTTPStatusError):
adapter.fetch_barcode("x")
assert calls["n"] == 1 # 404 is not retried
def test_retries_connection_error_then_succeeds():
calls = {"n": 0}
def handler(request: httpx.Request) -> httpx.Response:
calls["n"] += 1
if calls["n"] == 1:
raise httpx.ConnectError("boom")
return httpx.Response(200, json={"status": 1, "product": {"code": "y"}})
adapter = _adapter(handler, max_retries=4)
assert adapter.fetch_barcode("y") == {"code": "y"}
assert calls["n"] == 2
+11
View File
@@ -65,3 +65,14 @@ def test_transform_full_record():
def test_transform_drops_unnamed():
assert transform({"code": "0000000000000"}) is None
def test_transform_clamps_long_serving_size():
raw = {
"code": "3017624010701",
"product_name": "X",
"serving_size": "1 portion (30 g) / 1 portion (30 g) / extra long descriptive text here",
}
rec = transform(raw)
assert rec is not None
assert len(rec["food"]["serving_size"]) == 32
+1
View File
@@ -0,0 +1 @@
DROP TABLE IF EXISTS api_key;
+24
View File
@@ -0,0 +1,24 @@
-- API keys for the public read-only API. Keys grant higher rate limits and let
-- usage be attributed to a caller; the API itself stays free and read-only.
-- Only the SHA-256 hash of a key is stored; the plaintext is shown once at
-- creation time. Keys are issued/revoked from the admin console. The public
-- server only ever SELECTs from this table (request counting lives in Redis),
-- preserving its read-only contract against PostgreSQL.
CREATE TABLE IF NOT EXISTS api_key (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
name TEXT NOT NULL,
key_prefix VARCHAR(20) NOT NULL, -- shown for identification, e.g. og_live_AbC1
key_hash TEXT NOT NULL UNIQUE, -- hex SHA-256 of the full key
owner_email TEXT,
tier VARCHAR(16) NOT NULL DEFAULT 'free',
rate_limit_per_min INT NOT NULL DEFAULT 120,
revoked_at TIMESTAMPTZ,
created_by TEXT,
created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
CONSTRAINT api_key_tier_chk CHECK (tier IN ('free', 'partner', 'internal')),
CONSTRAINT api_key_rate_chk CHECK (rate_limit_per_min > 0)
);
-- Fast lookup of active keys by their hash on every authenticated request.
CREATE INDEX IF NOT EXISTS idx_api_key_active_hash
ON api_key (key_hash) WHERE revoked_at IS NULL;
+2
View File
@@ -0,0 +1,2 @@
DROP INDEX IF EXISTS idx_product_country;
DROP INDEX IF EXISTS idx_brand_name_trgm;
+9
View File
@@ -0,0 +1,9 @@
-- Search upgrade: trigram-based fuzzy matching + brand/country filters.
-- pg_trgm and the product.name GIN index already exist (see 0001). Add a
-- matching trigram index on brand.name so brand filtering / fuzzy brand
-- lookups can use an index instead of a sequential scan.
CREATE INDEX IF NOT EXISTS idx_brand_name_trgm ON brand USING gin (name gin_trgm_ops);
-- country_of_origin is filtered by exact/prefix match; a plain btree index
-- keeps that cheap as the catalog grows.
CREATE INDEX IF NOT EXISTS idx_product_country ON product (country_of_origin);
+15 -4
View File
@@ -25,11 +25,22 @@ export interface SearchResult {
total: number;
}
export interface SearchFilters {
brand?: string;
country?: string;
}
export const api = {
search: (q: string, page = 1, size = 20) =>
req<SearchResult>(
`/api/v1/products/search?q=${encodeURIComponent(q)}&page=${page}&size=${size}`,
),
search: (q: string, page = 1, size = 20, filters: SearchFilters = {}) => {
const params = new URLSearchParams({
q,
page: String(page),
size: String(size),
});
if (filters.brand) params.set("brand", filters.brand);
if (filters.country) params.set("country", filters.country);
return req<SearchResult>(`/api/v1/products/search?${params.toString()}`);
},
product: (id: string) => req<Product>(`/api/v1/products/${id}`),
categories: () => req<{ items: Category[] }>(`/api/v1/categories`),
submit: (input: SubmissionInput) =>
+58 -6
View File
@@ -117,7 +117,7 @@ export default function ApiDocs() {
<code className="font-mono bg-gray-100 rounded px-1.5 py-0.5">{BASE}</code>
</div>
<ul className="mt-2 list-disc pl-5 text-gray-600 space-y-1">
<li> API Key / Token GET </li>
<li> API Key / Token GET Key </li>
<li>
<code className="font-mono">page</code> 1
<code className="font-mono">size</code> 20 100
@@ -131,6 +131,53 @@ export default function ApiDocs() {
</div>
</div>
<div className="bg-white border rounded-lg p-5">
<h2 className="text-lg font-semibold text-gray-800"></h2>
<p className="mt-2 text-gray-600 text-sm leading-relaxed">
API <strong></strong> IP
API Key
</p>
<div className="mt-3">
<Code>{`# 二选一
curl -H "X-API-Key: og_live_xxxxxxxx" ${BASE}/products/search?q=牛奶
curl -H "Authorization: Bearer og_live_xxxxxxxx" ${BASE}/products/search?q=牛奶`}</Code>
</div>
<p className="mt-3 text-gray-600 text-sm leading-relaxed">
<strong></strong>便
</p>
<table className="mt-3 w-full text-sm">
<thead className="text-gray-400 text-left">
<tr>
<th className="font-medium pr-4 pb-1"></th>
<th className="font-medium pb-1"></th>
</tr>
</thead>
<tbody className="align-top">
<tr>
<td className="pr-4 py-0.5 font-mono text-gray-700">X-RateLimit-Limit</td>
<td className="py-0.5 text-gray-600"></td>
</tr>
<tr>
<td className="pr-4 py-0.5 font-mono text-gray-700">X-RateLimit-Remaining</td>
<td className="py-0.5 text-gray-600"></td>
</tr>
<tr>
<td className="pr-4 py-0.5 font-mono text-gray-700">X-RateLimit-Reset</td>
<td className="py-0.5 text-gray-600"> Unix </td>
</tr>
<tr>
<td className="pr-4 py-0.5 font-mono text-gray-700">Retry-After</td>
<td className="py-0.5 text-gray-600"></td>
</tr>
</tbody>
</table>
<p className="mt-3 text-gray-600 text-sm leading-relaxed">
<code className="font-mono">429 Too Many Requests</code>
<code className="font-mono">rate_limited</code> Key
<code className="font-mono">401</code> <code className="font-mono">invalid_api_key</code>
</p>
</div>
<Endpoint
method="GET"
path="/healthz"
@@ -164,19 +211,24 @@ export default function ApiDocs() {
method="GET"
path="/api/v1/products/search"
title="搜索商品"
desc="按名称模糊搜索,可按品类过滤,支持分页。"
desc="按名称做三元组(trigram)模糊搜索,可容忍错别字;支持品类/品牌/产地过滤;结果按相关度(名称相似度 × 数据质量分)排序。"
params={[
{ name: "q", desc: "关键词(名称/条码),留空返回全部" },
{ name: "category", desc: "品类编码过滤,如 food.beverages" },
{ name: "q", desc: "关键词(名称/条码),支持模糊匹配;留空则按质量分返回全部" },
{ name: "category", desc: "品类编码(含子树),如 food.beverages" },
{ name: "brand", desc: "品牌名(模糊匹配),如 Ferrero" },
{ name: "country", desc: "产地前缀(不区分大小写),如 China" },
{ name: "page", desc: "页码,默认 1" },
{ name: "size", desc: "每页条数,默认 20,最大 100" },
]}
example={`curl "${BASE}/products/search?q=nutella&page=1&size=20"`}
example={`curl "${BASE}/products/search?q=nutela&country=Italy&page=1&size=20"`}
response={`{
"items": [
{ "id": "…", "gtin": "3017624010701",
"name": "Nutella", "brand": "Ferrero",
"category_path": "food.snacks.chocolate" }
"category_path": "food.snacks.chocolate",
"country_of_origin": "Italy",
"quality_score": 0.81,
"score": 0.71 }
],
"page": 1, "size": 20, "total": 1
}`}
+26 -2
View File
@@ -13,6 +13,8 @@ export default function Home({
onApi: () => void;
}) {
const [q, setQ] = useState("");
const [brand, setBrand] = useState("");
const [country, setCountry] = useState("");
const [items, setItems] = useState<ProductSummary[]>([]);
const [total, setTotal] = useState(0);
const [searched, setSearched] = useState(false);
@@ -24,7 +26,10 @@ export default function Home({
setLoading(true);
setError("");
try {
const res = await api.search(q.trim(), 1, 30);
const res = await api.search(q.trim(), 1, 30, {
brand: brand.trim() || undefined,
country: country.trim() || undefined,
});
setItems(res.items);
setTotal(res.total);
setSearched(true);
@@ -61,6 +66,20 @@ export default function Home({
{loading ? "检索中…" : "检索"}
</button>
</form>
<div className="mt-3 max-w-2xl mx-auto flex flex-wrap items-center justify-center gap-2 text-sm">
<input
value={brand}
onChange={(e) => setBrand(e.target.value)}
placeholder="按品牌筛选(如 Ferrero"
className="flex-1 min-w-[14rem] bg-white border rounded-lg px-3 py-2 outline-none focus:ring-2 focus:ring-emerald-300"
/>
<input
value={country}
onChange={(e) => setCountry(e.target.value)}
placeholder="按产地筛选(如 China"
className="flex-1 min-w-[14rem] bg-white border rounded-lg px-3 py-2 outline-none focus:ring-2 focus:ring-emerald-300"
/>
</div>
<button
onClick={onApi}
className="mt-4 inline-flex items-center gap-1.5 text-sm text-emerald-700 hover:underline"
@@ -105,7 +124,12 @@ export default function Home({
{p.gtin ? ` · ${p.gtin}` : ""}
</div>
</div>
<span className="text-xs text-gray-400">{p.category_path || ""}</span>
<span className="text-xs text-gray-400 text-right">
<span className="block">{p.category_path || ""}</span>
{p.country_of_origin ? (
<span className="block text-gray-400">{p.country_of_origin}</span>
) : null}
</span>
</button>
</li>
))}
+3
View File
@@ -4,6 +4,9 @@ export interface ProductSummary {
name: string;
brand: string | null;
category_path: string | null;
country_of_origin?: string | null;
quality_score?: number;
score?: number | null;
}
export interface Barcode {
+9
View File
@@ -0,0 +1,9 @@
# build artifacts
bypos-collector.exe
bypos-collector
bypos-collector-linux
*.exe
# collected data / outputs
*.jsonl
*.csv
*.log
+67
View File
@@ -0,0 +1,67 @@
# bypos-collector
按条码批量采集商品档案的小工具(单文件 Windows/Linux 程序,自带本地 Web 控制台)。
数据来源是云店「新增商品」输入条码时所查的同一个中心商品库 `zc.bypos.net`
采集结果落地为 JSONL,供后续导入本项目(天工/goods)。
> 这是一个**独立模块**(有自己的 `go.mod`),与 `api/` 主服务互不影响,
> CI 不会编译它。放在 `tools/` 下仅作代码留存与后期迭代。
## 目录
| 文件 | 说明 |
| --- | --- |
| `main.go` | 入口:本地 HTTP 服务 + 启动浏览器 + API 路由(start/stop/stats/download/export.csv) |
| `collect.go` | 核心:签名、EAN-13 校验位、范围/清单枚举、并发+限速、JSONL 落库、断点续采 |
| `web/index.html` | 内嵌(`go:embed`)的控制台界面 |
| `build.sh` | 交叉编译出 `bypos-collector.exe`(windows/amd64)与 linux 测试二进制 |
| `使用说明.md` | 面向使用者的操作说明 |
## 构建
```bash
./build.sh
# 产物:bypos-collector.exe(发给 Windows 用户)/ bypos-collector-linux(本地测试)
```
二进制与采集产物(`*.jsonl`/`*.csv`)已在 `.gitignore` 中排除,不入库。
## 接口与签名(逆向所得,后期迭代参考)
云店新增商品页输入条码时,前端经服务端代理 `/prod-api/ZmSvr/httpUtil/getGet`
转发到中心库:
```
GET http://zc.bypos.net/byGoodsService/byMessage.asmx/GetGoodsinfo
?sdogid=<账号id>&regnum=1&barcode=<条码>
&sparm1=<md5(sdogid)> # 常量,随账号固定
&sparm2=<md5(barcode + tsMs)> # tsMs = 当前秒*1000(末尾恒为 000)
&sparm3=<tsMs 前 10 位 = 秒级时间戳>
&sparm4=&barcodetype=yunpos
```
返回 `<string>{...json...}</string>`,内层 JSON 字段:
| 上游字段 | 含义 | 归一化字段 |
| --- | --- | --- |
| item_name | 品名 | name |
| item_size | 规格 | spec |
| unit_no | 单位 | unit |
| item_area | 产地/地区 | area |
| birth_com | 生产企业(常空) | manufacturer |
| birth_doc | 生产许可(常空) | license |
| inprice | 建议进价 | in_price |
| sellprice | 建议零售价 | sell_price |
| retcode | 1=命中,0=失败 | status(hit/miss/invalid) |
`retmsg` 含「非国标条码 / 参数异常」=> invalid;含「条码不存在」=> miss。
## 配置
- `sdogid`:中心库账号 id(本项目所属云店账号的授权 id)。默认值见
`collect.go``defaultSdogID`,也可用 `-sdogid` 参数或控制台覆盖。
**这是账号级凭证**——若本仓库对外公开,建议改为从环境变量/外部配置读取。
## 注意
批量自动查询比页面逐条更"重",上游可能对账号限频。请低速、分前缀/品类分批采集。
+11
View File
@@ -0,0 +1,11 @@
#!/usr/bin/env bash
# Build the bypos-collector for Windows (and a Linux binary for local testing).
set -euo pipefail
cd "$(dirname "$0")"
gofmt -w ./*.go
go vet ./...
echo "building windows/amd64 .exe ..."
GOOS=windows GOARCH=amd64 go build -ldflags "-s -w" -o bypos-collector.exe .
echo "building linux/amd64 (for testing) ..."
go build -o bypos-collector-linux .
ls -la bypos-collector.exe bypos-collector-linux
+481
View File
@@ -0,0 +1,481 @@
package main
import (
"context"
"crypto/md5"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"net/http"
"net/url"
"os"
"path/filepath"
"regexp"
"strconv"
"strings"
"sync"
"sync/atomic"
"time"
)
// ---- upstream config ----
// sdogid is the 云店 account license id observed in the live request. It is
// configurable so the tool is not tied to a single account.
const defaultSdogID = "137966"
const endpoint = "http://zc.bypos.net/byGoodsService/byMessage.asmx/GetGoodsinfo"
var stringTagRe = regexp.MustCompile(`(?s)<string[^>]*>(.*)</string>`)
// Product is the normalized record we persist (one JSON object per line).
type Product struct {
Barcode string `json:"barcode"`
Name string `json:"name"` // item_name 品名
Spec string `json:"spec"` // item_size 规格
Unit string `json:"unit"` // unit_no 单位
Area string `json:"area"` // item_area 产地/地区
Manufacturer string `json:"manufacturer"` // birth_com 生产企业
License string `json:"license"` // birth_doc 生产许可
InPrice string `json:"in_price"` // 建议进价
SellPrice string `json:"sell_price"` // 建议零售价
Status string `json:"status"` // hit / miss / invalid / error
RetMsg string `json:"retmsg"` // 原始返回信息
FetchedAt string `json:"fetched_at"` // RFC3339
Source string `json:"source"` // zc.bypos.net
}
// upstream raw fields
type rawResp struct {
RetCode string `json:"retcode"`
RetMsg string `json:"retmsg"`
Barcode string `json:"barcode"`
ItemName string `json:"item_name"`
UnitNo string `json:"unit_no"`
ItemSize string `json:"item_size"`
ItemArea string `json:"item_area"`
BirthCom string `json:"birth_com"`
BirthDoc string `json:"birth_doc"`
InPrice string `json:"inprice"`
SellPrice string `json:"sellprice"`
}
func md5hex(s string) string {
h := md5.Sum([]byte(s))
return hex.EncodeToString(h[:])
}
// ean13Check computes the EAN-13 check digit for a 12-digit body.
func ean13Check(body string) (string, bool) {
if len(body) != 12 {
return "", false
}
sum := 0
for i := 0; i < 12; i++ {
c := body[i]
if c < '0' || c > '9' {
return "", false
}
d := int(c - '0')
if i%2 == 0 {
sum += d
} else {
sum += d * 3
}
}
chk := (10 - (sum % 10)) % 10
return body + strconv.Itoa(chk), true
}
// lookup queries the upstream central library for one barcode.
func (c *Collector) lookup(ctx context.Context, barcode string) (*Product, error) {
tsMs := strconv.FormatInt(time.Now().Unix()*1000, 10) // always ends in 000
sparm1 := md5hex(c.sdogID)
sparm2 := md5hex(barcode + tsMs)
sparm3 := tsMs[:10]
q := url.Values{}
q.Set("sdogid", c.sdogID)
q.Set("regnum", "1")
q.Set("barcode", barcode)
q.Set("sparm1", sparm1)
q.Set("sparm2", sparm2)
q.Set("sparm3", sparm3)
q.Set("sparm4", "")
q.Set("barcodetype", "yunpos")
reqURL := endpoint + "?" + q.Encode()
req, _ := http.NewRequestWithContext(ctx, "GET", reqURL, nil)
req.Header.Set("User-Agent", "Mozilla/5.0")
resp, err := c.client.Do(req)
if err != nil {
return nil, err
}
defer resp.Body.Close()
b, err := io.ReadAll(resp.Body)
if err != nil {
return nil, err
}
inner := b
if m := stringTagRe.FindSubmatch(b); m != nil {
inner = m[1]
}
var r rawResp
if err := json.Unmarshal(inner, &r); err != nil {
return nil, fmt.Errorf("parse: %v (body=%.120s)", err, string(b))
}
p := &Product{
Barcode: barcode,
RetMsg: r.RetMsg,
FetchedAt: time.Now().Format(time.RFC3339),
Source: "zc.bypos.net",
}
if r.RetCode == "1" {
p.Status = "hit"
p.Name = strings.TrimSpace(r.ItemName)
p.Spec = strings.TrimSpace(r.ItemSize)
p.Unit = strings.TrimSpace(r.UnitNo)
p.Area = strings.TrimSpace(r.ItemArea)
p.Manufacturer = strings.TrimSpace(r.BirthCom)
p.License = strings.TrimSpace(r.BirthDoc)
p.InPrice = strings.TrimSpace(r.InPrice)
p.SellPrice = strings.TrimSpace(r.SellPrice)
} else if strings.Contains(r.RetMsg, "非国标") || strings.Contains(r.RetMsg, "参数异常") {
p.Status = "invalid"
} else {
p.Status = "miss"
}
return p, nil
}
// ---- job / collector state ----
type Stats struct {
Running bool `json:"running"`
Total int64 `json:"total"`
Done int64 `json:"done"`
Hits int64 `json:"hits"`
Miss int64 `json:"miss"`
Invalid int64 `json:"invalid"`
Errors int64 `json:"errors"`
Skipped int64 `json:"skipped"`
Current string `json:"current"`
OutFile string `json:"out_file"`
StartedAt string `json:"started_at"`
Message string `json:"message"`
}
type Collector struct {
mu sync.Mutex
client *http.Client
sdogID string
outPath string
outFile *os.File
cancel context.CancelFunc
wg sync.WaitGroup
// atomic counters
total, done, hits, miss, invalid, errors, skipped int64
running int32
current atomic.Value // string
startedAt string
message string
seen map[string]struct{} // barcodes already in output (dedupe / resume)
recent []Product // ring of last results for UI
}
func NewCollector(sdogID string) *Collector {
if sdogID == "" {
sdogID = defaultSdogID
}
c := &Collector{
client: &http.Client{Timeout: 25 * time.Second},
sdogID: sdogID,
seen: map[string]struct{}{},
}
c.current.Store("")
return c
}
func (c *Collector) isRunning() bool { return atomic.LoadInt32(&c.running) == 1 }
func (c *Collector) snapshot() Stats {
cur, _ := c.current.Load().(string)
c.mu.Lock()
msg := c.message
out := c.outPath
started := c.startedAt
c.mu.Unlock()
return Stats{
Running: c.isRunning(),
Total: atomic.LoadInt64(&c.total),
Done: atomic.LoadInt64(&c.done),
Hits: atomic.LoadInt64(&c.hits),
Miss: atomic.LoadInt64(&c.miss),
Invalid: atomic.LoadInt64(&c.invalid),
Errors: atomic.LoadInt64(&c.errors),
Skipped: atomic.LoadInt64(&c.skipped),
Current: cur,
OutFile: out,
StartedAt: started,
Message: msg,
}
}
func (c *Collector) recentResults() []Product {
c.mu.Lock()
defer c.mu.Unlock()
out := make([]Product, len(c.recent))
copy(out, c.recent)
return out
}
func (c *Collector) pushRecent(p Product) {
c.mu.Lock()
c.recent = append(c.recent, p)
if len(c.recent) > 60 {
c.recent = c.recent[len(c.recent)-60:]
}
c.mu.Unlock()
}
// loadSeen reads an existing output file to build the dedupe set (for resume).
func (c *Collector) loadSeen(path string) error {
c.seen = map[string]struct{}{}
f, err := os.Open(path)
if err != nil {
if os.IsNotExist(err) {
return nil
}
return err
}
defer f.Close()
dec := json.NewDecoder(f)
for {
var p Product
if err := dec.Decode(&p); err != nil {
break
}
if p.Barcode != "" {
c.seen[p.Barcode] = struct{}{}
}
}
return nil
}
type JobReq struct {
Mode string `json:"mode"` // "range" | "list"
StartBody string `json:"start_body"` // 12-digit body (range mode)
EndBody string `json:"end_body"` // 12-digit body (range mode)
List string `json:"list"` // newline/space separated barcodes (list mode)
Concurrency int `json:"concurrency"` // parallel requests
DelayMs int `json:"delay_ms"` // min interval between request starts
LogMiss bool `json:"log_miss"` // also write miss/invalid lines
OutFile string `json:"out_file"`
SdogID string `json:"sdog_id"`
}
func sanitizeBarcodes(s string) []string {
fields := regexp.MustCompile(`[^0-9]+`).Split(s, -1)
var out []string
for _, f := range fields {
if f != "" {
out = append(out, f)
}
}
return out
}
// Start launches a collection job. Returns error if validation fails or busy.
func (c *Collector) Start(req JobReq) error {
if c.isRunning() {
return fmt.Errorf("已有任务在运行")
}
if req.Concurrency <= 0 {
req.Concurrency = 3
}
if req.Concurrency > 20 {
req.Concurrency = 20
}
if req.DelayMs < 0 {
req.DelayMs = 0
}
if req.OutFile == "" {
req.OutFile = "products.jsonl"
}
if req.SdogID != "" {
c.sdogID = req.SdogID
}
// Build the list of barcodes to query.
var barcodes []string
switch req.Mode {
case "list":
barcodes = sanitizeBarcodes(req.List)
if len(barcodes) == 0 {
return fmt.Errorf("条码清单为空")
}
case "range":
start, err := strconv.ParseInt(req.StartBody, 10, 64)
if err != nil || len(req.StartBody) != 12 {
return fmt.Errorf("起始码必须是 12 位数字(不含校验位)")
}
end, err := strconv.ParseInt(req.EndBody, 10, 64)
if err != nil || len(req.EndBody) != 12 {
return fmt.Errorf("结束码必须是 12 位数字(不含校验位)")
}
if end < start {
return fmt.Errorf("结束码不能小于起始码")
}
if end-start+1 > 5_000_000 {
return fmt.Errorf("单次范围过大(>500万),请缩小区间分批采集")
}
for v := start; v <= end; v++ {
body := fmt.Sprintf("%012d", v)
full, ok := ean13Check(body)
if ok {
barcodes = append(barcodes, full)
}
}
default:
return fmt.Errorf("未知模式: %s", req.Mode)
}
abs, _ := filepath.Abs(req.OutFile)
if err := c.loadSeen(abs); err != nil {
return fmt.Errorf("读取已有文件失败: %v", err)
}
f, err := os.OpenFile(abs, os.O_CREATE|os.O_WRONLY|os.O_APPEND, 0644)
if err != nil {
return fmt.Errorf("打开输出文件失败: %v", err)
}
c.outFile = f
c.outPath = abs
// reset counters
atomic.StoreInt64(&c.total, int64(len(barcodes)))
atomic.StoreInt64(&c.done, 0)
atomic.StoreInt64(&c.hits, 0)
atomic.StoreInt64(&c.miss, 0)
atomic.StoreInt64(&c.invalid, 0)
atomic.StoreInt64(&c.errors, 0)
atomic.StoreInt64(&c.skipped, 0)
c.mu.Lock()
c.recent = nil
c.startedAt = time.Now().Format(time.RFC3339)
c.message = ""
c.mu.Unlock()
atomic.StoreInt32(&c.running, 1)
ctx, cancel := context.WithCancel(context.Background())
c.cancel = cancel
go c.run(ctx, barcodes, req)
return nil
}
func (c *Collector) Stop() {
if c.cancel != nil {
c.cancel()
}
}
func (c *Collector) run(ctx context.Context, barcodes []string, req JobReq) {
defer func() {
atomic.StoreInt32(&c.running, 0)
if c.outFile != nil {
c.outFile.Sync()
c.outFile.Close()
c.outFile = nil
}
c.current.Store("")
}()
jobs := make(chan string, req.Concurrency*2)
var writeMu sync.Mutex
// global rate limiter: one token every DelayMs
var ticker *time.Ticker
if req.DelayMs > 0 {
ticker = time.NewTicker(time.Duration(req.DelayMs) * time.Millisecond)
defer ticker.Stop()
}
worker := func() {
defer c.wg.Done()
for bc := range jobs {
if ctx.Err() != nil {
return
}
if ticker != nil {
select {
case <-ticker.C:
case <-ctx.Done():
return
}
}
c.current.Store(bc)
p, err := c.lookup(ctx, bc)
if err != nil {
if ctx.Err() != nil {
return
}
atomic.AddInt64(&c.errors, 1)
atomic.AddInt64(&c.done, 1)
ep := Product{Barcode: bc, Status: "error", RetMsg: err.Error(), FetchedAt: time.Now().Format(time.RFC3339), Source: "zc.bypos.net"}
c.pushRecent(ep)
continue
}
switch p.Status {
case "hit":
atomic.AddInt64(&c.hits, 1)
case "miss":
atomic.AddInt64(&c.miss, 1)
case "invalid":
atomic.AddInt64(&c.invalid, 1)
}
atomic.AddInt64(&c.done, 1)
c.pushRecent(*p)
if p.Status == "hit" || req.LogMiss {
line, _ := json.Marshal(p)
writeMu.Lock()
c.outFile.Write(line)
c.outFile.Write([]byte("\n"))
c.seen[bc] = struct{}{}
writeMu.Unlock()
}
}
}
for i := 0; i < req.Concurrency; i++ {
c.wg.Add(1)
go worker()
}
for _, bc := range barcodes {
if ctx.Err() != nil {
break
}
if _, ok := c.seen[bc]; ok {
atomic.AddInt64(&c.skipped, 1)
atomic.AddInt64(&c.done, 1)
continue
}
select {
case jobs <- bc:
case <-ctx.Done():
}
}
close(jobs)
c.wg.Wait()
c.mu.Lock()
if ctx.Err() != nil {
c.message = "已停止"
} else {
c.message = "采集完成"
}
c.mu.Unlock()
}
+3
View File
@@ -0,0 +1,3 @@
module byposcollector
go 1.23.4
+156
View File
@@ -0,0 +1,156 @@
package main
import (
"embed"
"encoding/csv"
"encoding/json"
"flag"
"fmt"
"io"
"io/fs"
"log"
"net"
"net/http"
"os"
"os/exec"
"runtime"
"time"
)
//go:embed web/*
var webFS embed.FS
var collector = NewCollector("")
func main() {
addr := flag.String("addr", "127.0.0.1:8765", "本地监听地址")
noOpen := flag.Bool("no-open", false, "不自动打开浏览器")
sdog := flag.String("sdogid", "", "中心库账号 id(默认使用内置值)")
flag.Parse()
if *sdog != "" {
collector.sdogID = *sdog
}
sub, _ := fs.Sub(webFS, "web")
mux := http.NewServeMux()
mux.Handle("/", http.FileServer(http.FS(sub)))
mux.HandleFunc("/api/start", handleStart)
mux.HandleFunc("/api/stop", handleStop)
mux.HandleFunc("/api/stats", handleStats)
mux.HandleFunc("/api/download", handleDownload)
mux.HandleFunc("/api/export.csv", handleExportCSV)
ln, err := net.Listen("tcp", *addr)
if err != nil {
log.Fatalf("无法监听 %s: %v", *addr, err)
}
realAddr := ln.Addr().String()
urlStr := "http://" + realAddr + "/"
fmt.Println("==============================================")
fmt.Println(" 中心库商品采集器 bypos-collector")
fmt.Println(" 控制台: " + urlStr)
fmt.Println(" 关闭本窗口即停止程序")
fmt.Println("==============================================")
if !*noOpen {
go openBrowser(urlStr)
}
log.Fatal(http.Serve(ln, mux))
}
func openBrowser(url string) {
time.Sleep(600 * time.Millisecond)
var cmd string
var args []string
switch runtime.GOOS {
case "windows":
cmd = "rundll32"
args = []string{"url.dll,FileProtocolHandler", url}
case "darwin":
cmd = "open"
args = []string{url}
default:
cmd = "xdg-open"
args = []string{url}
}
_ = exec.Command(cmd, args...).Start()
}
func writeJSON(w http.ResponseWriter, code int, v interface{}) {
w.Header().Set("Content-Type", "application/json; charset=utf-8")
w.WriteHeader(code)
json.NewEncoder(w).Encode(v)
}
func handleStart(w http.ResponseWriter, r *http.Request) {
if r.Method != "POST" {
http.Error(w, "method", 405)
return
}
var req JobReq
if err := json.NewDecoder(r.Body).Decode(&req); err != nil {
writeJSON(w, 400, map[string]string{"error": "请求格式错误"})
return
}
if err := collector.Start(req); err != nil {
writeJSON(w, 400, map[string]string{"error": err.Error()})
return
}
writeJSON(w, 200, map[string]string{"ok": "started"})
}
func handleStop(w http.ResponseWriter, r *http.Request) {
collector.Stop()
writeJSON(w, 200, map[string]string{"ok": "stopping"})
}
func handleStats(w http.ResponseWriter, r *http.Request) {
writeJSON(w, 200, map[string]interface{}{
"stats": collector.snapshot(),
"recent": collector.recentResults(),
})
}
func handleDownload(w http.ResponseWriter, r *http.Request) {
s := collector.snapshot()
if s.OutFile == "" {
http.Error(w, "no output yet", 404)
return
}
f, err := os.Open(s.OutFile)
if err != nil {
http.Error(w, err.Error(), 404)
return
}
defer f.Close()
w.Header().Set("Content-Type", "application/x-ndjson; charset=utf-8")
w.Header().Set("Content-Disposition", "attachment; filename=products.jsonl")
io.Copy(w, f)
}
func handleExportCSV(w http.ResponseWriter, r *http.Request) {
s := collector.snapshot()
if s.OutFile == "" {
http.Error(w, "no output yet", 404)
return
}
f, err := os.Open(s.OutFile)
if err != nil {
http.Error(w, err.Error(), 404)
return
}
defer f.Close()
w.Header().Set("Content-Type", "text/csv; charset=utf-8")
w.Header().Set("Content-Disposition", "attachment; filename=products.csv")
w.Write([]byte{0xEF, 0xBB, 0xBF}) // UTF-8 BOM so Excel reads Chinese correctly
cw := csv.NewWriter(w)
cw.Write([]string{"barcode", "name", "spec", "unit", "area", "manufacturer", "license", "in_price", "sell_price", "status", "fetched_at"})
dec := json.NewDecoder(f)
for {
var p Product
if err := dec.Decode(&p); err != nil {
break
}
cw.Write([]string{p.Barcode, p.Name, p.Spec, p.Unit, p.Area, p.Manufacturer, p.License, p.InPrice, p.SellPrice, p.Status, p.FetchedAt})
}
cw.Flush()
}
+206
View File
@@ -0,0 +1,206 @@
<!DOCTYPE html>
<html lang="zh-CN">
<head>
<meta charset="utf-8"/>
<meta name="viewport" content="width=device-width, initial-scale=1"/>
<title>中心库商品采集器</title>
<style>
* { box-sizing: border-box; }
body { font-family: -apple-system, "Microsoft YaHei", Arial, sans-serif; margin: 0; background:#f4f6f9; color:#222; }
header { background:#1f6feb; color:#fff; padding:14px 22px; font-size:18px; font-weight:600; }
.wrap { max-width:1080px; margin:18px auto; padding:0 16px; }
.card { background:#fff; border:1px solid #e3e8ef; border-radius:10px; padding:18px 20px; margin-bottom:16px; }
.card h3 { margin:0 0 12px; font-size:15px; color:#1f6feb; }
label { display:block; font-size:13px; color:#555; margin:8px 0 4px; }
input[type=text], input[type=number], textarea, select {
width:100%; padding:8px 10px; border:1px solid #cdd5e0; border-radius:6px; font-size:14px;
}
textarea { height:90px; font-family:monospace; }
.row { display:flex; gap:14px; flex-wrap:wrap; }
.row > div { flex:1; min-width:160px; }
.tabs { display:flex; gap:8px; margin-bottom:12px; }
.tab { padding:7px 16px; border:1px solid #cdd5e0; border-radius:20px; cursor:pointer; font-size:13px; background:#fff; }
.tab.active { background:#1f6feb; color:#fff; border-color:#1f6feb; }
button.primary { background:#1f6feb; color:#fff; border:none; padding:10px 22px; border-radius:6px; font-size:14px; cursor:pointer; }
button.danger { background:#d1242f; color:#fff; border:none; padding:10px 22px; border-radius:6px; font-size:14px; cursor:pointer; }
button.ghost { background:#fff; color:#1f6feb; border:1px solid #1f6feb; padding:8px 16px; border-radius:6px; cursor:pointer; font-size:13px; }
button:disabled { opacity:.5; cursor:not-allowed; }
.stats { display:flex; gap:10px; flex-wrap:wrap; }
.stat { flex:1; min-width:90px; background:#f7f9fc; border:1px solid #e3e8ef; border-radius:8px; padding:10px; text-align:center; }
.stat .n { font-size:22px; font-weight:700; }
.stat .l { font-size:12px; color:#777; margin-top:2px; }
.bar { height:10px; background:#e3e8ef; border-radius:6px; overflow:hidden; margin:10px 0; }
.bar > div { height:100%; background:#2da44e; width:0%; transition:width .4s; }
table { width:100%; border-collapse:collapse; font-size:13px; }
th, td { text-align:left; padding:6px 8px; border-bottom:1px solid #eef1f5; white-space:nowrap; overflow:hidden; text-overflow:ellipsis; max-width:180px; }
th { color:#888; font-weight:600; }
.hit { color:#2da44e; } .miss { color:#999; } .invalid { color:#d1242f; } .error { color:#bf8700; }
.hint { font-size:12px; color:#888; margin-top:6px; line-height:1.5; }
.est { font-size:13px; color:#1f6feb; margin-top:6px; }
</style>
</head>
<body>
<header>中心库商品采集器 · bypos-collector</header>
<div class="wrap">
<div class="card">
<h3>① 规划采集范围</h3>
<div class="tabs">
<div class="tab active" data-mode="range" onclick="setMode('range')">按条码范围</div>
<div class="tab" data-mode="list" onclick="setMode('list')">按条码清单</div>
</div>
<div id="pane-range">
<div class="hint">EAN-13 国标条码共 13 位,最后一位是校验位由程序自动计算。下面填<b>前 12 位</b>(本体),程序逐个枚举并补校验位查询。常见前缀:69 开头为中国大陆。</div>
<label>快捷填充前缀(可选)</label>
<div class="row">
<div><input type="text" id="prefix" placeholder="如 690100,点下方按钮自动算区间"/></div>
<div style="flex:0"><button class="ghost" onclick="fillFromPrefix()">用前缀填充区间</button></div>
</div>
<div class="row">
<div>
<label>起始本体(12 位)</label>
<input type="text" id="start" value="690100000000" maxlength="12"/>
</div>
<div>
<label>结束本体(12 位)</label>
<input type="text" id="end" value="690100000999" maxlength="12"/>
</div>
</div>
<div class="est" id="est"></div>
</div>
<div id="pane-list" style="display:none">
<label>粘贴条码清单(每行一个,或用空格/逗号分隔)</label>
<textarea id="list" placeholder="6901028941068&#10;6920202888883"></textarea>
</div>
</div>
<div class="card">
<h3>② 采集参数</h3>
<div class="row">
<div>
<label>并发数</label>
<input type="number" id="concurrency" value="3" min="1" max="20"/>
</div>
<div>
<label>每次请求间隔(毫秒)</label>
<input type="number" id="delay" value="300" min="0"/>
</div>
<div>
<label>输出文件名</label>
<input type="text" id="outfile" value="products.jsonl"/>
</div>
</div>
<label style="margin-top:12px"><input type="checkbox" id="logmiss"/> 同时记录未命中/无效条码(默认只存命中)</label>
<div class="hint">速度越快越容易触发上游频控。建议并发 3、间隔 300ms 起步,稳定后再调。已采过的条码会自动跳过(断点续采)。</div>
</div>
<div class="card">
<h3>③ 运行</h3>
<div style="margin-bottom:12px">
<button class="primary" id="btnStart" onclick="start()">开始采集</button>
<button class="danger" id="btnStop" onclick="stop()" disabled>停止</button>
<button class="ghost" onclick="window.open('/api/download')">下载 JSONL</button>
<button class="ghost" onclick="window.open('/api/export.csv')">导出 CSV(Excel)</button>
</div>
<div class="bar"><div id="prog"></div></div>
<div class="stats">
<div class="stat"><div class="n" id="s-done">0</div><div class="l">已处理</div></div>
<div class="stat"><div class="n" id="s-total">0</div><div class="l">总计</div></div>
<div class="stat"><div class="n hit" id="s-hits">0</div><div class="l">命中</div></div>
<div class="stat"><div class="n miss" id="s-miss">0</div><div class="l">未命中</div></div>
<div class="stat"><div class="n invalid" id="s-invalid">0</div><div class="l">无效</div></div>
<div class="stat"><div class="n error" id="s-errors">0</div><div class="l">错误</div></div>
<div class="stat"><div class="n" id="s-skipped">0</div><div class="l">跳过</div></div>
</div>
<div class="hint" id="msg"></div>
</div>
<div class="card">
<h3>④ 实时结果(最近 60 条)</h3>
<div style="max-height:340px; overflow:auto">
<table>
<thead><tr><th>条码</th><th>品名</th><th>规格</th><th>单位</th><th>产地</th><th>进价</th><th>零售价</th><th>状态</th></tr></thead>
<tbody id="rows"></tbody>
</table>
</div>
</div>
</div>
<script>
let mode = 'range';
function setMode(m){
mode = m;
document.querySelectorAll('.tab').forEach(t=>t.classList.toggle('active', t.dataset.mode===m));
document.getElementById('pane-range').style.display = m==='range'?'block':'none';
document.getElementById('pane-list').style.display = m==='list'?'block':'none';
}
function fillFromPrefix(){
let p = document.getElementById('prefix').value.replace(/[^0-9]/g,'');
if(!p){ alert('请先填前缀'); return; }
if(p.length>=12){ alert('前缀太长,应少于 12 位'); return; }
let pad = 12 - p.length;
document.getElementById('start').value = p + '0'.repeat(pad);
document.getElementById('end').value = p + '9'.repeat(pad);
updateEst();
}
function updateEst(){
let s = document.getElementById('start').value.replace(/[^0-9]/g,'');
let e = document.getElementById('end').value.replace(/[^0-9]/g,'');
if(s.length===12 && e.length===12){
let n = (BigInt(e) - BigInt(s)) + 1n;
document.getElementById('est').textContent = '本次将查询约 ' + n.toString() + ' 个条码';
} else {
document.getElementById('est').textContent = '';
}
}
document.getElementById('start').addEventListener('input', updateEst);
document.getElementById('end').addEventListener('input', updateEst);
updateEst();
async function start(){
let body = {
mode: mode,
start_body: document.getElementById('start').value.trim(),
end_body: document.getElementById('end').value.trim(),
list: document.getElementById('list').value,
concurrency: parseInt(document.getElementById('concurrency').value)||3,
delay_ms: parseInt(document.getElementById('delay').value)||0,
log_miss: document.getElementById('logmiss').checked,
out_file: document.getElementById('outfile').value.trim()
};
let r = await fetch('/api/start', {method:'POST', headers:{'Content-Type':'application/json'}, body:JSON.stringify(body)});
let j = await r.json();
if(j.error){ alert('启动失败: ' + j.error); return; }
}
async function stop(){ await fetch('/api/stop', {method:'POST'}); }
function esc(s){ return (s||'').replace(/[&<>]/g, c=>({'&':'&amp;','<':'&lt;','>':'&gt;'}[c])); }
async function poll(){
try{
let r = await fetch('/api/stats'); let j = await r.json();
let s = j.stats;
document.getElementById('s-done').textContent = s.done;
document.getElementById('s-total').textContent = s.total;
document.getElementById('s-hits').textContent = s.hits;
document.getElementById('s-miss').textContent = s.miss;
document.getElementById('s-invalid').textContent = s.invalid;
document.getElementById('s-errors').textContent = s.errors;
document.getElementById('s-skipped').textContent = s.skipped;
let pct = s.total>0 ? Math.floor(s.done*100/s.total) : 0;
document.getElementById('prog').style.width = pct + '%';
document.getElementById('msg').textContent = (s.running? ('采集中… 当前 '+s.current) : (s.message||'空闲'));
document.getElementById('btnStart').disabled = s.running;
document.getElementById('btnStop').disabled = !s.running;
let rows = (j.recent||[]).slice().reverse().map(p=>
'<tr><td>'+esc(p.barcode)+'</td><td>'+esc(p.name)+'</td><td>'+esc(p.spec)+'</td><td>'+esc(p.unit)+'</td><td>'+esc(p.area)+'</td><td>'+esc(p.in_price)+'</td><td>'+esc(p.sell_price)+'</td><td class="'+p.status+'">'+esc(p.status)+'</td></tr>'
).join('');
document.getElementById('rows').innerHTML = rows;
}catch(e){}
}
setInterval(poll, 1000); poll();
</script>
</body>
</html>
+68
View File
@@ -0,0 +1,68 @@
# 中心库商品采集器 bypos-collector 使用说明
一个单文件 Windows 小程序,通过云店「新增商品」用到的同一个中心商品库
(`zc.bypos.net`)按条码批量采集商品档案(品名/规格/单位/产地/厂商/建议进价/建议零售价),
存到本地,供后续导入天工(goods)系统。
## 一、运行
1.`bypos-collector.exe` 放到任意空文件夹(采集结果会生成在同一文件夹)。
2. 双击运行。会弹出一个黑色命令行窗口(不要关它),并自动打开浏览器控制台
`http://127.0.0.1:8765/`
- 若没自动打开,手动在浏览器输入上面这个地址。
3. 用完直接关掉那个命令行窗口即可退出。
## 二、采集
控制台分四步:
**① 规划采集范围** —— 两种方式二选一:
- **按条码范围**:EAN-13 国标条码共 13 位,最后一位是校验位,程序自动算。
你只填**前 12 位**的起止区间。可在「快捷填充前缀」里填如 `690100`,
点按钮自动生成区间(`690100000000` ~ `690100999999`)。
- **按条码清单**:直接粘贴一批条码(每行一个,或空格/逗号分隔)。
**② 采集参数**:
- 并发数(默认 3)、请求间隔(默认 300ms):**越慢越安全**,上游可能对账号限频。
- 输出文件名(默认 `products.jsonl`)。
- 「同时记录未命中/无效条码」:默认只存命中的;勾上会把未命中也记下来。
**③ 运行**:点「开始采集」。进度、命中/未命中/错误实时显示。
已经采过的条码会自动跳过(可随时停了再开,断点续采)。
**④ 实时结果**:最近 60 条滚动显示。
## 三、导出
- 「下载 JSONL」:原始数据(每行一个 JSON),用于导入天工系统。
- 「导出 CSV」:Excel 可直接打开查看。
## 四、字段说明(JSONL 每行)
| 字段 | 含义 |
| --- | --- |
| barcode | 条码(GTIN/EAN-13) |
| name | 品名 |
| spec | 规格 |
| unit | 单位 |
| area | 产地/地区 |
| manufacturer | 生产企业(常为空) |
| license | 生产许可(常为空) |
| in_price | 建议进价 |
| sell_price | 建议零售价 |
| status | hit=命中 / miss=不存在 / invalid=非国标条码 / error=请求出错 |
| fetched_at | 采集时间 |
## 五、注意
- 这是用云店账号授权去查上游中心库,**批量自动**比页面里一条条查更"重",
上游厂商可能对账号做频控/限额。请低速、分批("一点点采"),发现大量报错就降速。
- 全量 69 段是个天文数字,不要无脑全跑;建议按你关心的品牌/品类前缀分批。
## 六、命令行参数(可选)
```
bypos-collector.exe -addr 127.0.0.1:8765 # 改监听端口
bypos-collector.exe -no-open # 不自动开浏览器
bypos-collector.exe -sdogid 137966 # 指定中心库账号 id
```