52 Commits

Author SHA1 Message Date
Misaka
bccf7cd396 fix(runtime): strip host-injected env vars when launching 安能 2026-08-02 14:14:34 +08:00
Misaka
c1bd53d832 feat(tasks): record trigger mode, target date and force in task history 2026-08-02 14:14:30 +08:00
Misaka
3c32720985 feat(sites): auto-screenshot on final download failure for debugging
所有站点 with_retry 在最后一次重试失败、reset 之前自动截图,
保存到 logs/screenshots/。网页站点走 Playwright page.screenshot(),
安能走 CDP Page.captureScreenshot。截图失败绝不阻塞任务流程。

- paths.py: 新增 SCREENSHOT_DIR (BASE_DIR/logs/screenshots/)
- runtime.py: 新增 capture_error_screenshot() 工具函数
- .gitignore: 新增 logs/ 忽略规则

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-02 11:18:21 +08:00
Misaka
53c71aeeac refactor(status): derive site ready/business_date from PG instead of ingest_state
_ready_flags 改从 PostgreSQL 直接查询(expected_record /
actual_record / baishi_daily_stats),target_date = today − offset。
消除因 ingested_at 日期比对导致的每日零点全站 ready 集体重置。

store.py: 新增 has_data(site, kind, target_date) 查 PG 数据存在性
runtime.py: _ready_flags 返回 (flags, dates) 同源元组,_apply_ready
  同步写入 business_date,修正 ready 与 business_date 不同源导致的
  前端日期标签漂移(如实到就绪却显示'前天')

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-02 09:22:54 +08:00
Misaka
837264f7b0 fix(zto): resolve real-today date picker failure on month-boundary days
The ZTO jQuery Date Range Picker always displays a dual-month view.
When today falls on the 1st (or early days) of a month, the same date
appears in both panels: a hidden ghost cell (month1, display:none) and
a visible cell (month2). Both carry the real-today CSS class, so .first
picks the hidden one, causing wait_for(visible) to timeout.

Replace DOM-based real-today time extraction with Python datetime
computation. Add _zto_find_visible_day to locate the actually visible
date cell (skipping hidden ghost cells, trying both midnight and
23:59:59 time variants). Fix _zto_flip_to_target_month to use the
same visibility-aware lookup.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-01 17:50:30 +08:00
Misaka
eaf56c1c9e refactor(status): derive readiness from ingest_state; drop Excel probe & /data
状态盘就绪态 {kind}_ready 改为从持久的 ingest_state 派生(DB 真相、重启不丢),
不再由心跳读 downloads/ Excel、也不在启动时重置:
- 心跳 _fresh_ingest/_ready_flags/_apply_ready:expected/actual_ready=该类今天入库成功;
  undelivered_ready=百世原生(今天入库) / 4 站派生(expected ∧ actual)。
- _persist_to_db 入库后调 _refresh_ready 立即派生(省 30s 心跳等待,与心跳同源);
  4 站 undelivered 连入 expected+actual,故 ingest_state 补记 expected/actual/undelivered 三行。
- _record_business_date 只写 business_date,不再碰 ready。
- 删 run_heartbeat 的 Excel 探测循环、DATA_FILENAMES、probe_data_file、遗物清理。
- state_store 增 set_business_date / set_ready(仅写单字段,不碰彼此)。
- 删 GET /data/{filename}(前端不再下载原始中转 Excel);/report 保留。

原则:Excel 只作「站点下载→入库」中转,状态/比对一律走 DB。
真机+单元验证:重启后心跳派生(百世今日入库→立即绿、重启不掉灰);韵达 expected-only
时 undelivered 不亮、补 actual 后亮(派生);百世原生;_ready_flags 7 例边界全过。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 23:47:23 +08:00
Misaka
dc7653c256 fix(db_compare): anchor comparison on actual_offset so yunda reads fresh data
比对以实到扫描日(actual.scan_time::date)为锚,锚点日期必须等于实到下载日 = actual_offset。
原 _site_undelivered_handler(runtime.py)与 _target_date_for(db_compare.py)误用
expected_offset 算锚点:韵达 exp=1 / act=0,锚点落到 today-1(昨天),读到历史数据。
两处改用 actual_offset 后韵达锚点 = today,反推出 expected 的昨天批次正确参与比对。
其余 3 站 exp == act == 0,锚点不变。

5 站真机验证:韵达 07-31 比对现读新鲜数据 107/104/3(修复前读昨天 116/115/1);
中通/安能/顺心 不变;全站汇总报表重新生成。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 22:58:12 +08:00
Misaka
3c7e9f2522 feat(db_compare): DB-based full summary report + wire 跑比对 to it
新增 build_full_report:4 站走 compare_site_date、百世走 baishi_daily_stats(基数) + undelivered_record(按 ingested_at 日期过滤),复用 compare.build_summary 渲染 KPI/柱状图/口径说明,产 output/应到未到数据.xlsx。跑比对入口(__compare__)从 compare.main() Excel 路径切换到 build_full_report。百世未到件用基数差(undelivered_pieces)与应到/已到自洽。

集成验证:前端跑比对 -> build_full_report -> /report 下载,各站未到件 顺心8/中通10/韵达2/安能0/百世8。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 21:56:39 +08:00
Misaka
8521c200ab feat(store): persist 百世 daily basis to PG (baishi_daily_stats)
百世应到/实到基数(应扫/已扫)原仅在 state_store(单值、无历史)。新增百世专用聚合表 baishi_daily_stats(site+business_date UPSERT),baishi 下载时抓到基数直接落库(store.upsert_baishi_daily_stats,一步,不绕 state_store→store)。state_store 双写保留以兼容旧 Excel 汇总(process_baishi),后续统一清理。

真机端到端验证通过:PG (2026-07-31, 194, 186, 8),state_store 一致。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 21:23:08 +08:00
Misaka
c6ad6a0ca2 fix(runtime): ingest before db_compare so first-run/force reads fresh data
DB 比对发生在入库之前,导致首次/force 时 PG 无当天数据,比对返回 None、不产出 Excel。在 _site_undelivered_handler 下载成功后、比对前,前置 _record_business_date + ingest_task,使比对能读到本次下载的数据。dispatch_task 后置 _persist_to_db 保持不变(对 undelivered 幂等重复一次,安全)。

四站真机验证通过(中通/韵达/安能 + 顺心):前置入库日志均出现在 db_compare 之前,force 首跑即产出 Excel。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 20:25:52 +08:00
Misaka_Company
541836fd1b feat(db_compare): add PostgreSQL-based comparison engine with SF handling
Replace Excel-based undelivered comparison with DB queries for all four
sites. The engine anchors on actual scan_time, reverse-lookups handover
batches, and compares expected vs actual waybill-by-waybill.

Shunxin SF waybills: use COUNT(*) instead of COUNT(DISTINCT piece_no)
since SF piece numbers are random and not derivable from the waybill.

Changes:
- db_compare.py: new module with compare_site_date(), compare_site_batch(),
  write_result_excel(), and POST /compare API endpoint
- runtime.py: switch _site_undelivered_handler from compare.write_site_file
  (Excel) to db_compare (DB); downloads succeed independently of comparison
- server.py: add POST /compare endpoint with date validation
- docs: implementation plan for Shunxin DB comparison

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 16:31:54 +08:00
Misaka_Company
95597fbb0c docs: add four-site comparison logic review report
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 14:48:24 +08:00
Misaka_Company
f66e6dd39e fix(yunda): invert handover number filter to keep empty rows
The previous filter kept rows with non-empty handover numbers
(派件/签收 scans), which were duplicate rows. The correct logic
is to keep rows with empty handover numbers (到/接件 scans).

- store.py: change != "" to == "" in ingest filter
- compare.py: add same filter before comparison (previously missing)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-31 14:44:50 +08:00
Misaka
09e05f8dfc feat(runtime): disable microphone by default to suppress permission popup
韵达等站点打开时会请求麦克风权限,触发浏览器系统级授权弹窗。在共享 context 上加
add_init_script,于每个页面/iframe 加载前覆盖 navigator.mediaDevices.getUserMedia
(及 webkit/moz 旧版) 为直接 reject(NotAllowedError: Permission disabled):站点调用时
立即被拒、不再弹出系统授权窗,且麦克风被真正挡住(非授权给它)。物流工作台无需音视频
采集,故对所有网页站点统一生效。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 20:58:46 +08:00
Misaka
047036f46b feat(yunda): cross-month calendar navigation for date-specific download
韵达两个流程的日期组件不同,分别适配跨月翻月:
- 应到(.layui-laydate 新版):读 .laydate-set-ym 当前年月,点 .laydate-prev-m/.laydate-next-m
  翻月,格子 td[lay-ymd='YYYY-M-D']。
- 实到(#laydate_box 旧版):读 #laydate_y/#laydate_m 输入框值,点 #laydate_MM 内
  .laydate_chprev/.laydate_chnext 翻月,格子 td[y][m][d]。
两处设日期改走 _yunda_pick_laydate_new / _yunda_pick_laydate_old,同月(offset 当月)
行为不变,仅跨月时翻月。年*12+月 比较天然支持跨年。

实到 #startDate/#endDate 保持普通 click(force=True 会在导航后 laydate 绑定完成前
抢先点击导致面板打不开);新增 #laydate_box:visible 就绪等待。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 20:26:45 +08:00
Misaka
21662d3944 feat(shunxin): cross-month calendar navigation for date-specific download
顺心 Ant Design 日期面板跨月导航:目标日期不在当前月视窗时,读面板头部
.ant-picker-year-btn/.ant-picker-month-btn 得当前年月,按差值点
.ant-picker-header-prev-btn/next-btn 翻到目标月再选格子。应到(车辆点到)与
实到(卸车扫描记录)两处设日期统一改走 _shunxin_pick_date。

同月(offset 当月)行为不变,仅在跨月时多走翻月;与中通 _zto_flip_to_target_month
思路对称,适配 Ant Design 面板。年*12+月 比较天然支持跨年。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 19:58:10 +08:00
Misaka_Company
4f0ef69739 fix(server): lower date backtrack limit from 90 to 31 days
Aneng only supports querying data up to 31 days back. The previous
90-day cap allowed dates that would silently fail at download time,
so restrict the POST /tasks `date` validation to a 31-day window.
2026-07-29 16:55:14 +08:00
Misaka_Company
5e2ef72afd feat(server): add date field to POST /tasks with legality validation
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:11:12 +08:00
Misaka_Company
fcb75643d0 feat(runtime): propagate date through dispatch chain and business-date snapshot
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:10:23 +08:00
Misaka_Company
cda764a305 feat(anneng): support date arg (date takes precedence over offset)
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:09:30 +08:00
Misaka_Company
e51aab2e0b feat(shunxin): support date arg, propagate to both accounts
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:08:51 +08:00
Misaka_Company
bacccc43ab feat(yunda): support date arg (date takes precedence over offset)
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:08:06 +08:00
Misaka_Company
e240d92ad9 feat(zto): support date arg via effective-offset (reuses cross-month nav)
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:07:27 +08:00
Misaka_Company
30e9203fca feat(baishi): accept date kwarg (ignored) for unified dispatch signature
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:06:37 +08:00
Misaka_Company
93e48118bf docs: add implementation plan for date-specific download API
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 13:03:49 +08:00
Misaka_Company
a68fd46c51 docs: add design spec for date-specific download API (developer interface)
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 12:58:08 +08:00
Misaka_Company
e2ae3acd68 feat(zto): cross-month calendar navigation instead of fallback to today
When the offset target date falls outside the current two-month view, flip .prev months (1 month/step) until the target day cell enters month1 view, instead of silently falling back to today. JS-dispatched clicks avoid the .date-range-length-tip hover interception.

Also switch the expected (#beginDate) day-cell click to _dom_click (matching actual #daterange): the range-length tooltip intercepts the second click on non-today cells during single-day range selection, causing timeouts.

Verified end-to-end: offset=45 (-> 2026-06-14, cross-month) expected task succeeded and downloaded 6-14 data without fallback.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 12:49:44 +08:00
Misaka_Company
f710622a3d feat: dedup expected data by handover_no before export submit
- store.get_existing_handover_nos: query PG expected_record.handover_no

- inject dedup skip before submitting export in zto/yunda/anneng/shunxin

- shunxin reads RTS handover_no from waybill-list view (method 1)

- force-redownload switch threaded via task_spec -> dispatch -> impl

- schema: add idx_expected_handover index

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-29 10:57:07 +08:00
Misaka
d74a35844f feat(schedule): 周期抓取调度取代每日定点下载
应到/实到数据各自独立周期抓取(启用+激活时段+频率),抓取与处理解耦,抓完自动落库(复用现有 _persist_to_db 钩子)。

- state_store: 新增 fetch_schedule 表(site,kind);所有连接加 timeout=3.0 防并发锁;get/set/get_all_fetch_schedules、create_task_if_idle(周期去重)、allowed_kinds;get_all_config 改返回 fetch_schedules;旧 set_schedule 标废弃。修复 get_all_fetch_schedules 列错位(_fetch_spec(r[2:]))。
- server: IntervalTrigger + _in_active_window(含跨午夜) + _enqueue_fetch(就绪门/激活窗口/去重) + _reschedule_fetch + lifespan 按(site,kind)注册 + FetchScheduleSpec + PUT/GET /config 新结构。
- runtime/store/BFF: 零改动。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-28 22:58:47 +08:00
Misaka
82ab6d8be3 feat(runtime): 服务模式任务执行不再激活浏览器窗口
网页站点(顺心/百世/中通/韵达)执行下载流程时不再把浏览器窗口置顶,避免打扰用户当前工作;流程在后台静默跑。启动登录阶段与交互模式不受影响。

- RuntimeContext 加 foreground 标志:True=任务执行时置顶(交互调试),False=后台静默(服务模式)
- launch_and_prepare(foreground=True) 透传;server worker 以 foreground=False 启动
- _web_handler 按 ctx.foreground 决定单 page 置顶;顺心(list) 透传给 shunxin_download
- shunxin 两个 download 加 foreground 参数,逐账号 bring_to_front 加条件
- 启动登录/初始弹窗清理的 bring_to_front 不变(启动时窗口需可见)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-27 23:17:10 +08:00
Misaka
29edcbfdbb fix(yunda): 导出面板适配 Element UI 改版
韵达导出面板从 jQuery(.allRight/#submitbutton) 改版为 Element UI,应到/实到导出均卡在全选字段步骤。

- actual/expected 改 Element UI 交互:全选 → 向右转移(el-icon-d-arrow-right) → 导出(el-icon-download) → 等 el-loading-mask → el-message-box 成功提示 → 确定
- 新增 _resolve_export_frame() 探测导出 iframe(实到=myFrame、应到=target1),都未命中时 dump 面板内 iframe 名便于排查
- 外层 layui-layer 弹层、查询、切导出服务下载逻辑不变

验证:韵达应到 49 条 / 实到 141 条下载成功。

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-27 22:50:47 +08:00
Misaka_Company
5e4889e845 docs: sync docs with current package layout and auto-ingest hook
- config.example.yaml: replace stale site_yunda/main_router with new
  module paths (sites.yunda, runtime)
- CLAUDE.md: add ingest-one to DB CLI list; new subsection documenting
  the download->PostgreSQL auto-ingest hook (_persist_to_db, ingest_task,
  ingest_state, /status.ingest, auto_ingest config)
- README.md: store tree comment lists all 5 CLI commands (add ingest-one)
- docs: /status row notes the ingest field; anneng CDP guide snippet
  gets timeout=15
- cli/server.py: docstring run command -> python -m inbound_verify.cli.server

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 12:35:54 +08:00
Misaka_Company
0111734e9c fix(anneng): bound CDP websocket recv with 15s socket timeout
websocket.create_connection had no timeout, so a hung Runtime.evaluate
(Electron business tab not replying) blocked ws.recv() indefinitely — the
export-poll's 300s deadline and with_retry could never fire (observed as a
~22min hang on an 安能 expected download). A 15s socket timeout lets recv
raise WebSocketTimeoutException so wait_until/with_retry can fail and retry.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 12:22:22 +08:00
Misaka_Company
615661c275 feat(server): expose ingest state in /status; doc auto-ingest in README 2026-07-24 11:16:19 +08:00
Misaka_Company
327f77a727 fix(runtime): wrap ingest_enabled gate so _persist_to_db never raises
Move the store.ingest_enabled() gate inside the main try/except. Previously
it sat outside as a standalone block: ingest_enabled() -> _load_pg_config()
raises FileNotFoundError when config.yaml is absent/malformed, which escaped
_persist_to_db, was caught by dispatch_task's outer except, and flipped a
successful download to FAILED -- violating the hook's never-raise invariant.
Now config errors print a [warn], record ok=False, and return silently.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 11:10:45 +08:00
Misaka_Company
62a66467a9 feat(runtime): auto-ingest hook after successful download
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 11:04:04 +08:00
Misaka_Company
7211846375 feat(state_store): add ingest_state table + set/get helpers 2026-07-24 10:55:43 +08:00
Misaka_Company
1093256de6 fix(store): guard ingest_task against 百世 non-undelivered kinds 2026-07-24 10:46:51 +08:00
Misaka_Company
0c51a41cfc feat(store): add kind-level ingest_task + ingest-one CLI subcommand
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 10:40:51 +08:00
Misaka_Company
70f518a1c3 feat(store): add auto_ingest config + pg connect/statement timeouts
Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 10:34:27 +08:00
Misaka_Company
81dea38310 docs: add ingest-hook design spec and implementation plan
Spec for mounting PostgreSQL ingest as a best-effort, kind-level,
synchronous hook on runtime.dispatch_task after a successful download,
plus the 5-task implementation plan.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 10:27:00 +08:00
Misaka_Company
f61f52bcbe docs: align docs with post-restructure package layout
Update README tree (compare.py + domain.py), CLAUDE.md module refs, and the three docs/ deep-dives to the current inbound_verify package (compare/domain/sites/cli). Correct the statistics review report's methodology to the current 口径 (应到=交接件数, 实到=直接数单号去重, 未到=应到−实到) per compare.py docstring; remove ghost-script references in the CDP guide.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 10:12:40 +08:00
Misaka_Company
79c84a5a0c fix: retry launch page.goto to absorb transient DNS/timeout
Third launch failure this session (after two 顺心 load-timeouts and a 中通 ERR_NAME_NOT_RESOLVED) — each a transient network blip that killed the whole worker because launch_and_prepare opened sites with no retry. Wrap each site open+goto in _open_page: up to 3 attempts (2s apart) on domcontentloaded, so transient DNS/timeout is absorbed; only persistent failure (all 3) aborts. All 5 sites + 安能 now launch clean.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 09:06:02 +08:00
Misaka_Company
62c44b262a fix: launch page.goto uses domcontentloaded to avoid slow-site timeout
launch_and_prepare opened each site with Playwright's default wait_until='load', which waits for every resource (ads/trackers/images). sxne.sxjdfreight.com (顺心) reliably takes >30s to fire load, so its goto timed out and — launch not being retried — killed the whole worker, leaving only one page open. Switch the two launch gotos to wait_until='domcontentloaded' (return as soon as the DOM is ready; the readiness polling still gates login). Behavior for already-fast sites unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 08:59:56 +08:00
Misaka_Company
23042c272b refactor: rename expected_undelivered to compare
git mv expected_undelivered.py -> compare.py; update the 3 importers (store/runtime/router) to import compare. All public names (main, write_site_file, _read_business_dates) unchanged. Also add the Tier 2 plan doc.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 08:40:04 +08:00
Misaka_Company
91ef8970b8 refactor: extract domain.py (shared site/file/colmap config)
Move ALL_REPORT_SITES / SITE_UNDELIVERED_FILE / BAISHI_FILE / BAISHI_COLUMNS / arrived_pieces_* / STATIONS / _site_cfg out of expected_undelivered into a new leaf module inbound_verify/domain.py. store.py and runtime.py now read site/file config from domain directly instead of through the compare engine (store keeps eu only for _read_business_dates). Behavior identical.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 08:38:18 +08:00
Misaka_Company
bd0df8909f refactor: expected_undelivered uses paths.DOWNLOAD_DIR/OUTPUT_DIR (Tier 2)
Remove the duplicate BASE/DOWNLOADS/OUTPUT self-anchor (the one that caused the Tier 1 hotfix bug when the file moved into the package). DOWNLOADS/OUTPUT now come from paths.py; OUTFILE derives from OUTPUT_DIR. Behavior identical.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-24 08:34:57 +08:00
Misaka_Company
17293bee79 fix: correct expected_undelivered.BASE anchor after package move
Tier 1 regression: expected_undelivered.py carries its own BASE=dirname(__file__) anchor (a duplicate of paths.py). The package move dropped the file one level deeper, so BASE resolved to inbound_verify/ and DOWNLOADS/OUTPUT pointed at non-existent inbound_verify/downloads|output — while site downloads write to the project-root dirs via paths.py. process()/write_site_file() thus skipped the compare with '[跳过] downloads 下缺少 ...', returned None/False, and every undelivered task for the 4 web/app sites reported failed (重试耗尽) despite the files downloading fine. Fix: anchor BASE two levels up (mirrors paths.py). Verified: write_site_file returns True; all 4 undelivered tasks now success with 未到 files generated.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-23 13:55:48 +08:00
Misaka_Company
5ac618c041 docs: update stale module-name references in comments and docstrings
Replace pre-rename references (main_router, db_store, server.py) in code comments and docstrings with their new locations (cli/router, store, cli/server, runtime). Comment-only; no behavior change.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-23 12:26:34 +08:00
Misaka_Company
a07b9b435b docs: update README/CLAUDE.md for package layout; add Tier 1 spec and plan
Rewrite run commands to python -m inbound_verify.* (and console_script aliases); add pip install -e . to env prep; refresh the directory tree. Also commit the design spec and Tier 1 implementation plan under docs/superpowers/.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-23 12:21:45 +08:00
Misaka_Company
7065b88269 refactor: move flat modules into inbound_verify package (Tier 1, behavior-identical)
Relocate 12 root .py modules into inbound_verify/ (sites/, cli/ subpackages). Rewrite all internal imports to package-qualified; drop the site_ prefix on the 5 site modules and their 28 call sites. Fix paths.py BASE_DIR to anchor at the project root. Add main() entry wrappers (cli/router, cli/server, store). No behavior change.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-23 12:21:31 +08:00
Misaka_Company
0174e64a04 chore: scaffold inbound_verify package and pyproject
Empty package markers (inbound_verify/, sites/, cli/) plus pyproject.toml with console_scripts entry points and dependency floors.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-07-23 12:20:54 +08:00
38 changed files with 6896 additions and 957 deletions

1
.gitignore vendored
View File

@@ -42,6 +42,7 @@ desktop.ini
downloads/
output/
state/
logs/
*.xlsx
*.xls
*.log

View File

@@ -8,73 +8,82 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
异常运单并汇总成 Excel。4 个网页站点 + 1 个 Electron 应用(安能)。
**两种运行模式**阶段1 起):
- **交互模式** `main_router.py`:人工调试 / 操作,交互菜单(登录、触发下载、[12] 状态盘)。
- **服务模式** `server.py`:常驻 + FastAPI客户端经 HTTP 触发任务、查状态、下载数据API 文档 `/docs`)。
- 两者共享 `runtime.py`(启动 / 就绪 / 任务派发 / 心跳)与 `state_store.py`SQLite 状态持久化)。
- **交互模式** `inbound_verify.cli.router`:人工调试 / 操作,交互菜单(登录、触发下载、[12] 状态盘)。
- **服务模式** `inbound_verify.cli.server`:常驻 + FastAPI客户端经 HTTP 触发任务、查状态、下载数据API 文档 `/docs`)。
- 两者共享 `inbound_verify.runtime`(启动 / 就绪 / 任务派发 / 心跳)与 `inbound_verify.state_store`SQLite 状态持久化)。
## 常用命令
所有 Python 一律在项目虚拟环境 `.venv` 中运行Windows 下可直接用
`.venv/Scripts/python.exe`,无需激活)。
`.venv/Scripts/python.exe`,无需激活)。首次 / 拉取新代码后需
`.venv/Scripts/python.exe -m pip install -e .`(以可编辑模式注册
`inbound-verify` 等命令)。
```bash
# 安装依赖(含安能 CDP 驱动所需的 websocket-client
# 安装依赖(含安能 CDP 驱动所需的 websocket-client+ 以可编辑模式注册命令
pip install -r requirements.txt
pip install -e .
playwright install chromium
# 运行主程序(交互式菜单,详见 main_router 的 run_multi_site_daemon
.venv/Scripts/python.exe main_router.py
# 运行主程序(交互式菜单,详见 inbound_verify.cli.router 的 run_multi_site_daemon
.venv/Scripts/python.exe -m inbound_verify.cli.router
# 装包后也可直接用命令inbound-verify
# 服务模式(常驻 + FastAPI客户端经 HTTP 触发;默认 :8000API 文档见 /docs
.venv/Scripts/python.exe server.py
.venv/Scripts/python.exe -m inbound_verify.cli.server
# 或inbound-verify-server
# DB CLI建库 / 初始化 / 灌数据 / 全流程 / 单站单类
.venv/Scripts/python.exe -m inbound_verify.store createdb # 或 init | ingest | ingest-one <site> <kind> | all
# 或inbound-verify-db createdb|init|ingest|ingest-one|all
# 单站点联调:在 config.yaml 设 debug.enabled=true + debug.target_site=顺心|百世|中通|韵达|安能
# 网页站:只挂载该站;安能:只启动 Electron 应用。
# 安能独立运行(需先以 --remote-debugging-port=9222 启动「安能全网门户.exe」并手动登录
.venv/Scripts/python.exe site_anneng.py expected # 或 actual
.venv/Scripts/python.exe -m inbound_verify.sites.anneng expected # 或 actual
# 格式化(全局规范:改完 Python 必须 Black
.venv/Scripts/python.exe -m black <file.py>
.venv/Scripts/python.exe -m black inbound_verify
# 语法自检
.venv/Scripts/python.exe -m py_compile <file.py>
.venv/Scripts/python.exe -m py_compile inbound_verify
```
**没有 pytest 测试套件。** "测试"指 `main_router` 菜单 **[8] 自动化测试**
**没有 pytest 测试套件。** "测试"指 `inbound_verify.cli.router` 菜单 **[8] 自动化测试**
`run_automation_test`,按 `CROSS_TEST_SEQUENCE` 交叉跑通各站点流程)。
## 架构big picture
### 运行模式与共享核心阶段0/1 重构)
- **`runtime.py`**:两种模式共享的核心——`launch_and_prepare`(启动 Playwright + 各站就绪 + 弹窗 + 心跳初值,阻塞至就绪)、`dispatch_task(ctx, {site,kind})`(派发任务,掉登录直接判 failed`run_heartbeat``probe_site_login/probe_data_file``RuntimeContext.stop()`。常量 `SITES_CONFIG/READY_SELECTORS/APP_SITES/HEARTBEAT_INTERVAL/DATA_FILENAMES` 在此。
- **`state_store.py`**SQLite 状态持久化(`state/state.db`)。`site_status`(登录态 + 数据态 + 时间戳,心跳刷新)、`task_history`(任务记录)。重启不丢。
- **`main_router.py`**交互模式菜单循环input 后台线程 + `_await_command` + `dispatch_task` + 心跳)。
- **`server.py`**:服务模式。**FastAPI主线程+ Playwright worker独立线程**——主线程处理 HTTP绝不碰 Playwrightworker 独占 page 操作,经 `task_queue` + `state_store` 通信。API`POST/GET /tasks``GET /status``GET /data/{file}`
- **`inbound_verify.runtime`**:两种模式共享的核心——`launch_and_prepare`(启动 Playwright + 各站就绪 + 弹窗 + 心跳初值,阻塞至就绪)、`dispatch_task(ctx, {site,kind})`(派发任务,掉登录直接判 failed`run_heartbeat``probe_site_login/probe_data_file``RuntimeContext.stop()`。常量 `SITES_CONFIG/READY_SELECTORS/APP_SITES/HEARTBEAT_INTERVAL/DATA_FILENAMES` 在此。
- **`inbound_verify.state_store`**SQLite 状态持久化(`state/state.db`)。`site_status`(登录态 + 数据态 + 时间戳,心跳刷新)、`task_history`(任务记录)。重启不丢。
- **`inbound_verify.cli.router`**交互模式菜单循环input 后台线程 + `_await_command` + `dispatch_task` + 心跳)。
- **`inbound_verify.cli.server`**:服务模式。**FastAPI主线程+ Playwright worker独立线程**——主线程处理 HTTP绝不碰 Playwrightworker 独占 page 操作,经 `task_queue` + `state_store` 通信。API`POST/GET /tasks``GET /status``GET /data/{file}`
- **关键线程约束**Playwright sync 对象绑定创建它的线程;`launch_and_prepare`(含 `sync_playwright().start()`)必须在持有 Playwright 的线程调用(交互=主线程,服务=worker 线程。FastAPI 路由绝不访问 page。
- 站点模块 `site_*.py``with_retry` 返回 `True/False`(成功 / 放弃),供 `dispatch_task` 判成败。
- 站点模块 `inbound_verify.sites.*``with_retry` 返回 `True/False`(成功 / 放弃),供 `dispatch_task` 判成败。
### 两套驱动模态 —— 这是理解全局的关键
- **网页 4 站**(顺心/百世/中通/韵达):`main_router` 用 Playwright 开 chromium
- **网页 4 站**(顺心/百世/中通/韵达):`inbound_verify.cli.router` 用 Playwright 开 chromium
每站一个 `page`,流程函数签名为 `xxx_download_impl(page)`。**例外:顺心是双账号**
——同一窗口开两个标签页(两个归属地账号),`pages_map["顺心"]` 存为 page **列表**
`shunxin_download(pages)` 接收列表(详见下文「顺心双账号」)。
- **安能**Electron 桌面应用,**不走 Playwright**。`main_router`
- **安能**Electron 桌面应用,**不走 Playwright**。`inbound_verify.cli.router`
`--remote-debugging-port=<动态空闲端口>` 启动 exe`launch_anneng`
通过 `site_anneng.set_cdp_port` 告知模块;`site_anneng.py` 用裸 CDPwebsocket
通过 `inbound_verify.sites.anneng.set_cdp_port` 告知模块;`inbound_verify.sites.anneng` 用裸 CDPwebsocket
驱动,业务 tab 是独立 webContents。这也是 `playwright-cli` 接管不了安能的原因
Electron 19 / Chrome 102 不支持 Playwright 要的 setDownloadBehavior
### 分层:路由纯调度,站点模块自洽
- `main_router.py` **只调度**:启动浏览器/安能、就绪轮询、登录检测、菜单分发。
菜单项直接调 `site_xxx.xxx_download(page)`(顺心传 page 列表),**不关心**重试/重置。
- 每个 `site_xxx.py` 对外只暴露"把任务做了"的入口,内部自洽:
- `inbound_verify.cli.router` **只调度**:启动浏览器/安能、就绪轮询、登录检测、菜单分发。
菜单项直接调 `sites.xxx_download(page)`(顺心传 page 列表),**不关心**重试/重置。
- 每个 `inbound_verify.sites.*` 模块对外只暴露"把任务做了"的入口,内部自洽:
- `xxx_download(...)` —— 公开入口,= `with_retry(站点, 标签, xxx_download_impl, xxx_reset)`
- `xxx_download_impl(...)` —— 单次执行、**无重试**(自动化测试刻意调它以探测原始失败)
- `xxx_reset(...)` —— 重置回初始态(网页 = `page.goto(HOME_URL)`;安能 = 关业务 tab + 收菜单)
- `with_retry(...)` —— 重试逻辑**内联在每个站点模块**(不抽公共组件,现阶段刻意不优化结构);
失败→重置→重试,最多 3 次(含首次),每次失败都重置(含最终放弃那次清场)
- `HOME_URL` —— 站点首页 URL`main_router.SITES_CONFIG` 引用它(单一来源)
- `HOME_URL` —— 站点首页 URL`runtime.SITES_CONFIG` 引用它(单一来源)
### 导出任务队列模式(顺心/中通/韵达/安能-应到 共用)
这些站的下载是异步的:提交导出(记 `export_times` 时间戳)→ 跳"导出任务管理"页轮询 →
@@ -88,28 +97,40 @@ playwright install chromium
### 顺心双账号(双归属地)
顺心业务上要同时处理**两个归属地网点**(两个账号)。程序在同一窗口开两个标签页,
人工分别登录两个账号(顺心站点支持同浏览器双账号并存,无需独立 context/窗口)。
- `main_router` 启动时为顺心开 2 个 `context.new_page()``pages_map["顺心"]` 为列表;
- `runtime` 启动时为顺心开 2 个 `context.new_page()``pages_map["顺心"]` 为列表;
就绪轮询要求**两个标签页都进主页**才算就绪;初始弹窗对两个标签页各处理一遍。
- `shunxin_expected_download(pages)` / `shunxin_actual_download(pages)` 接收 page 列表:
先用 `shunxin_belonging(page)` 读各账号归属地(首页「切换网点」控件 `.site___3o7nH`
**去重校验**(两账号同归属地则报错中止,防数据翻倍),再顺序对各账号跑一遍
`xxx_download_impl(page, out_tag=归属地)`(产物 `顺心-{归属}-{应到/实到}货物数据.xlsx`
最后 `shunxin_merge_final` 把两份 `pd.concat` 成统一的 `顺心-{应到/实到}货物数据.xlsx`
并删中间文件。比对层 `expected_undelivered` **零改动**(仍读同名文件)。
并删中间文件。比对层 `compare` **零改动**(仍读同名文件)。
- 导出队列不串扰:双账号**顺序执行**,账号 A 走完完整下载流程(远超 40s后 B 才提交,
配合每账号独立 `export_times` + ≤40s 容差B 不会误匹配 A 的任务。
### 比对
`expected_undelivered.py`(菜单 [9])纯离线:读 `downloads/` 下各站应到/实到 xlsx
`inbound_verify.compare`(菜单 [9])纯离线:读 `downloads/` 下各站应到/实到 xlsx
比对生成 `output/应到未到数据.xlsx`(汇总 + 各站明细)。
### 自动入库(下载成功后 → PostgreSQL
`dispatch_task` 下载成功后,在 `_record_business_date` 旁挂一个**尽力而为**钩子
`_persist_to_db(site, kind)``runtime`):懒导入 `store`,调 `store.ingest_task(site, kind)`
按 kind 幂等 UPSERT 进 PostgreSQLexpected/actual 各入其列undelivered——百世入未到、
4 站连入 expected+actual。**绝不影响下载任务的成功判定**:所有写库/写状态都包 try/except
失败只告警。
- 开关 `postgres.auto_ingest`(默认开)+ `connect_timeout_seconds`(兜 cpolar 抖动)。
- 可见性:结果写 `state_store.ingest_state`(每站每类 ok/count/ingested_at/error
`GET /status``ingest` 字段暴露。
- 手动:`store` CLI `ingest-one <site> <kind>` 用同套路由单测(脱离下载)。
- 注意:跑在 Playwright 线程、同步阻塞cpolar 慢/断靠超时 + try/except 降级,不重试不补入。
### 路径
`paths.py``DOWNLOAD_DIR` / `CONFIG_PATH` / `BASE_DIR` 全部锚定到项目目录,
`inbound_verify.paths``DOWNLOAD_DIR` / `CONFIG_PATH` / `BASE_DIR` 全部锚定到项目目录,
**不依赖运行时 cwd**——别用相对路径或 `os.getcwd()`
## 重要约定 / 易踩坑
- **登录是手动的**`main_router` 启动后会停在就绪轮询(`READY_SELECTORS` / `anneng_ready`
- **登录是手动的**`inbound_verify.cli.router` 启动后会停在就绪轮询(`READY_SELECTORS` / `anneng_ready`
直到检测到所有站点进入工作台才进菜单。仅韵达支持凭 `config.yaml` 凭据自动登录。
**顺心需登录两个账号**:同一窗口的两个标签页分别登录两个不同归属地账号,两个标签页
都进主页后才算就绪(顺心站点支持同浏览器双账号并存,故用同 context 标签页而非独立窗口)。

View File

@@ -32,19 +32,30 @@
```
InboundVerify/
├── main_router.py # 主入口 / 调度层(菜单、启动浏览器与安能、就绪轮询、登录检测
├── site_shunxin.py # 顺心站点模块(流程 + 重置 + 重试,自洽)
├── site_baishi.py # 百世站点模块
├── site_zto.py # 中通站点模块
├── site_yunda.py # 韵达站点模块(含自动登录
├── site_anneng.py # 安能站点模块Electron + CDP 驱动
├── expected_undelivered.py # 全站点应到未到离线比对,输出 output/应到未到数据.xlsx
├── paths.py # 统一路径锚点(以本目录为基准,不依赖 cwd
├── pyproject.toml # 打包 + 依赖 + console_scriptsinbound-verify 等
├── inbound_verify/ # 源码包
│ ├── paths.py # 统一路径锚点(以项目目录为基准,不依赖 cwd
│ ├── runtime.py # 两种模式共享核心(启动 / 就绪 / 任务派发 / 心跳)
├── state_store.py # SQLite 状态持久化state/state.db
│ ├── domain.py # 站点 / 文件名 / 列映射共享配置单一来源leaf
│ ├── compare.py # 全站点应到未到离线比对,输出 output/应到未到数据.xlsx
│ ├── store.py # DB CLI 入口createdb|init|ingest|ingest-one|all
│ ├── sites/ # 各站点模块(流程 + 重置 + 重试,自洽)
│ │ ├── shunxin.py # 顺心(含双账号)
│ │ ├── baishi.py # 百世
│ │ ├── zto.py # 中通
│ │ ├── yunda.py # 韵达(含自动登录)
│ │ └── anneng.py # 安能Electron + CDP 驱动)
│ └── cli/ # 命令行入口
│ ├── router.py # 交互菜单(调度层:启动 / 就绪轮询 / 登录检测 / 菜单分发)
│ └── server.py # FastAPI 服务模式(常驻 + HTTP 触发)
├── config.example.yaml # 配置模板
├── config.yaml # 真实配置(自行创建,已被 .gitignore 忽略)
├── requirements.txt
├── schema.sql # 数据库表结构store.py createdb / init 使用)
├── requirements.txt # pyproject 依赖的静态镜像
├── downloads/ # 各站点下载的原始数据
├── output/ # 比对报表输出
├── state/ # 运行状态持久化state.db
└── docs/ # 说明文档
```
@@ -63,10 +74,13 @@ python -m venv .venv
# 2. 安装依赖(含安能 CDP 驱动所需的 websocket-client
pip install -r requirements.txt
# 3. 安装 Playwright 浏览器内核(网页站点用
# 3. 以可编辑模式安装本包(注册 inbound-verify 等命令
pip install -e .
# 4. 安装 Playwright 浏览器内核(网页站点用)
playwright install chromium
# 4. 由模板创建本地配置并填入真实凭据
# 5. 由模板创建本地配置并填入真实凭据
cp config.example.yaml config.yaml
```
@@ -96,7 +110,10 @@ cp config.example.yaml config.yaml
## 五、运行
```bash
python main_router.py
# 交互菜单(任选其一)
python -m inbound_verify.cli.router
# 或装包后直接用命令:
inbound-verify
```
程序会:
@@ -120,16 +137,24 @@ python main_router.py
[0] 退出
```
### DB CLI入库
```bash
# DB CLI建库 / 初始化 / 灌数据 / 全流程 / 单站单类
.venv/Scripts/python.exe -m inbound_verify.store createdb # 或 init | ingest | ingest-one <site> <kind> | all
# 注下载成功后会自动入库postgres.auto_ingest默认开ingest-one 用于手动重灌指定站/类。
```
---
## 六、架构
**分层原则:路由层只调度,站点模块自洽。**
- **`main_router.py`(调度层)**:负责启动浏览器 / 安能、就绪轮询、登录检测、
- **`inbound_verify.cli.router`(调度层)**:负责启动浏览器 / 安能、就绪轮询、登录检测、
菜单分发。**不关心**"任务能否完成、失败怎么办"——只调
`site_xxx.xxx_download(page)` 然后等结果。
- **各 `site_xxx.py`(站点模块)**:每个模块对外只暴露一个"把任务做了"的入口
`inbound_verify.sites.xxx_download(page)` 然后等结果。
- **各 `inbound_verify.sites.*`(站点模块)**:每个模块对外只暴露一个"把任务做了"的入口
`xxx_download(...)`,内部自行处理一切:
- `HOME_URL`:站点首页 URL也供路由层 `SITES_CONFIG` 引用,单一来源);
- `xxx_reset(...)`:异常兜底的重置(网页 = 跳首页 URL安能 = 关业务 tab + 收菜单);
@@ -162,4 +187,4 @@ python main_router.py
- 首次运行需手动登录各站点(程序会停在就绪轮询,直到检测到所有站点进入工作台);
- 安能为单实例 Electron 应用:启动前请先关闭已打开的安能窗口;
- 所有下载/输出路径以项目目录为基准(见 `paths.py`),与从哪个目录启动无关。
- 所有下载/输出路径以项目目录为基准(见 `inbound_verify.paths`),与从哪个目录启动无关。

View File

@@ -46,7 +46,7 @@ zto:
# 韵达快运 (https://ky-sso.yunda56.com)
# ----------------------------------------------------------------------------
yunda:
# 自动登录的账号与密码(由 site_yunda.yunda_login 读取并填入登录表单)。
# 自动登录的账号与密码(由 inbound_verify.sites.yunda.yunda_login 读取并填入登录表单)。
# 留空时自动登录会填入空串导致登录失败,届时可在浏览器中改为手动登录。
# 真实凭据仅写进被 .gitignore 忽略的 config.yaml切勿提交示例值。
username: "YOUR_USERNAME_HERE"
@@ -58,9 +58,9 @@ yunda:
# ----------------------------------------------------------------------------
# 安能全网门户Electron 桌面应用,非网页)
# ----------------------------------------------------------------------------
# 与其他站点不同:安能不由 main_router 用浏览器打开,而是以调试模式启动其
# 与其他站点不同:安能不由浏览器打开,而是以调试模式启动其
# Electron 可执行文件(自动选取一个空闲端口作为 --remote-debugging-port避免端口冲突
# 启动后请在应用内手动登录,main_router 会自动轮询判断是否进入主页。
# 启动后请在应用内手动登录,runtime 会自动轮询判断是否进入主页。
anneng:
# 【已废弃】下载日期改由 Web 前端/API 按站点设偏移0=今天1=昨天…,存 state.db此项不再生效。
query_days: 1
@@ -70,14 +70,14 @@ anneng:
app_path: 'D:\SoftWare\SoftWare Installation\@ane-electron-uiapp\安能全网门户.exe'
# ----------------------------------------------------------------------------
# PostgreSQL 数据持久化(到货核销数据入库,详见 db_store.py
# PostgreSQL 数据持久化(到货核销数据入库,详见 store.py
# ----------------------------------------------------------------------------
# createdb 会连接名为 postgres 的维护库来创建下方 dbname 指定的数据库。
# 命令行:
# python db_store.py createdb 创建数据库(幂等)
# python db_store.py init 建表(幂等)
# python db_store.py ingest [site] 入库全站或单站(幂等 UPSERT
# python db_store.py all createdb → init → 全站 ingest
# python -m inbound_verify.store createdb 创建数据库(幂等)
# python -m inbound_verify.store init 建表(幂等)
# python -m inbound_verify.store ingest [site] 入库全站或单站(幂等 UPSERT
# python -m inbound_verify.store all createdb → init → 全站 ingest
postgres:
host: 127.0.0.1
port: 5432
@@ -87,3 +87,7 @@ postgres:
dbname: CQHXDB
# 承载到货核销表的专用 schema隔离 public表建在此 schema 下。
schema: inbound_verify
# 下载成功后自动入库(钩子,见 runtime._persist_to_dbfalse=跳过(无 PG/cpolar 的开发机)。
auto_ingest: true
# PG 连接超时cpolar 抖动时快速失败,不拖垮下载 worker。
connect_timeout_seconds: 5

View File

@@ -0,0 +1,102 @@
# 应到数据「提交导出任务前」去重 — 实现总结
> 日期2026-07-29
> 范围:顺心 / 中通 / 韵达 / 安能 4 站**应到expected**数据;百世与实到不在本次范围。
## 一、运行机制
在周期 / 手动触发下载时,于**提交导出任务之前**按交接单号判断该批应到数据是否已落库,已落库则跳过,从源头消除重复下载与重复落库。
```mermaid
flowchart TD
TRIG[周期调度 / 手动触发<br/>task_spec: site, kind, force] --> DISP[dispatch_task → 站点 download_impl]
DISP --> LOAD{force 强制重下?}
LOAD -- 是 --> EMPTY[existing = 空集]
LOAD -- 否 --> QRY[查 PG expected_record.handover_no]
QRY -- cpolar 失败 --> EMPTY
QRY -- 成功 --> SET[existing = 已落库交接单号集合]
EMPTY --> LOOP[遍历本次查询到的班次/交接单号]
SET --> LOOP
LOOP --> JUDGE{交接单号 ∈ existing?}
JUDGE -- 是 → 已落库 --> SKIP[⏭️ 跳过:不提交导出<br/>不 append export_times]
JUDGE -- 否 → 新单 --> EXP[提交导出任务 → 轮询下载 → 入库 UPSERT]
SKIP --> DONE{全部处理完}
EXP --> DONE
DONE --> FINAL{本次提交了新任务?}
FINAL -- 无 → 全跳过 --> BAIL[空兜底 return不进下载轮询]
FINAL -- 有 --> POLL[轮询导出任务管理页 → 下载 → 入库]
```
**核心要点:**
- **去重数据源**PostgreSQL `expected_record.handover_no`(已落库的权威记录),新增 `store.get_existing_handover_nos(site)` 查询。
- **判断时机**:提交导出任务**之前**(循环内逐单判断),而非下载之后。
- **安全降级**PG 不可用 / `force=true``existing=空集` → 当作未落库 → 继续提交(**宁可重复、绝不漏**UPSERT 兜底)。
- **空兜底**:全部跳过时 `export_times` 为空 → 直接 `return`,不进下载轮询(避免下载数校验失败 / 空转超时)。
- **force 开关**:前端 checkbox默认关`POST /tasks.force` → 一路透传到 impl周期调度恒不 force。
## 二、force 强制重下透传链路
```mermaid
flowchart LR
UI[前端 checkbox<br/>forceRedownload] --> POST["POST /api/tasks<br/>{site,kind,force}"]
POST --> BFF[Next BFF 透传]
BFF --> TS["task_spec<br/>{site,kind,force}"]
TS --> DISP[dispatch_task]
DISP --> HDR["handler(ctx, force)"]
HDR --> DL["download(pg, force)"]
DL --> IMPL["impl(pg, force)"]
IMPL --> DEC{force?}
DEC -- 是 --> EMPTY2["existing = 空集<br/>强制重下,跳过去重"]
DEC -- 否 --> LOAD2[查 PG 加载 existing]
```
> 周期调度(`_enqueue_fetch`)投递任务时不带 `force` → 默认不强制。
## 三、4 站点标识获取
| 站点 | 提交前标识 | 来源 |
| --- | --- | --- |
| 中通 / 韵达 / 安能 | 交接单号(原有代码已读取) | DOM 列 / CDP 复选框 |
| 顺心 | 交接单号 `RTS\d{3}WJ\d+` | 点"运单列表"后从界面读取方式1 |
> 顺心"班次号"业务上等同交接单号4 站统一用交接单号(= DB `handover_no`)作去重键。
## 四、改动概览
**后端 InboundVerify**
| 文件 | 改动 |
| --- | --- |
| `schema.sql` | +`idx_expected_handover` 索引 |
| `inbound_verify/store.py` | +`get_existing_handover_nos(site)`(含 cpolar 降级) |
| `inbound_verify/cli/server.py` | `TaskRequest.force` + 透传到 task_spec |
| `inbound_verify/runtime.py` | `dispatch_task` + 所有 handler 透传 `force``download(impl)` |
| `inbound_verify/sites/{zto,yunda,anneng,shunxin}.py` | `impl``force` + 提交导出前注入去重 + 空兜底 |
| `inbound_verify/sites/baishi.py` | `force` 形参兼容 |
**前端 dashboard**
| 文件 | 改动 |
| --- | --- |
| `app/page.tsx` | `forceRedownload` state + checkbox + `trigger`/`triggerPrimary`/`onTrigger` 透传 force |
> BFF `app/api/tasks/route.ts` 是 generic 透传,无需改动。
## 五、验证结论
4 站去重 + force 开关均实测通过:
| 站点 | 二次触发行为 | 结果 |
| --- | --- | --- |
| 中通 | `...801` 已入库 → 跳过 → 空兜底 return | 20svs 首次 59s ✅ |
| 顺心 | 双账号 4 班次全跳过(`RTS023WJ375478` 等) | ✅ |
| 韵达 | `...82001` 已入库 → 跳过 → 原有空兜底 | ✅ |
| 安能 | `4008242619171180544` 已入库 → 跳过 | **5svs 4 分钟)** ✅ |
| force | `[去重] 强制重下,跳过去重` + 已入库的重新导出 | ✅ |
实施过程中的两个问题均已解决:
1. **顺心 RTS 正则**`RTS\d+` 遇字母 W 停(只抓 `RTS023`)→ 改 `RTS[A-Z0-9]+` 抓完整 `RTS023WJ375320`
2. **韵达 force 偶发失败**:韵达站点自身 UI 不稳定(`section iframe` 匹配到 2 个 + 弹窗遮挡),与去重/force 无关;换安能验证 force 成功。
静态检查Blackpy310+ compileall + tsc 全绿。

View File

@@ -0,0 +1,399 @@
# 应到数据「提交导出任务前」去重 实施计划 v2
> **执行约定:** 本仓库无 pytest验证靠实跑站点流程 + DB 核对(见第 7 节)。
> 本仓库约定 **不自动提交**;所有改动落地后等用户明确说"提交"再 commit/push。
> 本次会话额外约定:未获用户明确指示前不动代码、不提交。
**Goal** 在周期性自动落库场景下,于「提交导出任务」之前,按**交接单号**(顺心=运单列表界面里的交接单号)判断该批应到数据是否已落库,已落库则跳过提交导出任务,从源头消除重复下载与重复落库;并提供一个"强制重下"开关(默认关)兜底。
**Architecture**
- 去重数据源 = PostgreSQL `expected_record.handover_no`(已存在字段,权威);`store.py` 新增 `get_existing_handover_nos(site)`(含 cpolar 降级)。
- 在 4 站**应到**下载循环内、提交导出动作之前注入"命中已落库则 `continue`"。
- `force` 开关经 `task_spec``dispatch_task` → handler → 各站 `download(page, force)``impl(page, force)` 透传;`force=True` 时跳过去重。周期 job 默认不 force。
**Tech Stack** Python 3.10+ / Playwright网页 3 站)/ 裸 CDP安能/ psycopg / SQLite前端 Next.js 16 + React 19。
## Global Constraints
- Python 一律 `.venv`;改完任何 `.py` 必须跑 `.venv/Scripts/python.exe -m black inbound_verify`
- 改完跑 `.venv/Scripts/python.exe -m py_compile inbound_verify` 自检。
- **不自动提交/推送**(覆盖全局 auto-push 默认)。
- **不动 `export_times` 时间容差≤40s匹配机制**CLAUDE.md 约定)。
- 安能 CDP 驱动,**绝不用 `Page.reload`**。
- `config.yaml` 已 gitignore不提交真实凭据。
- 前端是 **Next.js 16有 breaking changes**,写前端代码前先查 `node_modules/next/dist/docs/`
- 行号基于 2026-07-29 快照,实现时以当前代码为准、就近定位。
## 1. 已定决策v1 审核反馈)
| 决策 | 结论 |
|---|---|
| A. 去重数据源 | **查 PostgreSQL `expected_record.handover_no`** |
| B. 范围 | **本次只做应到expected**;实到/百世不动 |
| C. 强制开关 | **加 force 开关,默认不强制重下** |
| 顺心标识 | **方式1点"运单列表"后、点导出前,从运单列表界面读交接单号**(与其他 3 站统一用交接单号去重) |
## 2. 背景与问题根源
周期链路:`fetch_schedule(IntervalTrigger) → task_queue → worker → dispatch_task → handler → 提交导出+下载 → _persist_to_db(UPSERT)`
- DB 已幂等(`expected_record``(site, waybill_no)` UPSERT
-`dispatch_task` 调 handler 前**无"是否需要下载"判断**,周期触发重复"提交导出→下载→解析"。
- 本方案在「提交导出任务」前按交接单号去重,从源头省掉重复下载。
## 3. 各站探索结论(注入点)
| 站点 | 文件 | 提交导出位置 | 提交前标识 | 来源 |
|---|---|---|---|---|
| 中通 | `sites/zto.py` | `zto_expected_download_impl` L264 | ✅ 已有 `handover_no`L251 | 主表行 `td.nth(3)` 正则 18 位 |
| 韵达 | `sites/yunda.py` | `yunda_expected_download_impl` L342 | ✅ 已有 `raw_no`L291 | 列表行 `td.nth(1)` |
| 安能 | `sites/anneng.py` | 主循环 L927逐条 | ✅ 已有 `ewbs_no`L907-913 | CDP 复选框 `ewbsListNo=` 正则 19 位 |
| 顺心 | `sites/shunxin.py` | L319 点导出 | 🆕 方式1L316 后读运单列表界面交接单号 | DOM 待实勘 |
**顺心方式1关键事实**已探明raw 里「班次号」「交接单号」都有,且一个交接单 = 一个班次 = 多条运单;交接单号即入库 `handover_no`,与另 3 站同键。
## 4. 文件结构
| 文件 | 改动 |
|---|---|
| `schema.sql` | 加 `expected_record(site, handover_no)` 索引 |
| `inbound_verify/store.py` | 新增 `get_existing_handover_nos(site)` |
| `inbound_verify/cli/server.py` | `TaskRequest``force``create_task` 透传 force周期不 force |
| `inbound_verify/runtime.py` | `dispatch_task` 读 force 传 handler所有 handler 加 `force` 形参并透传到 `download_func` |
| `inbound_verify/sites/zto.py` | `download/impl``force`;应到循环注入去重 + 空兜底 |
| `inbound_verify/sites/yunda.py` | 同上(空兜底已存在) |
| `inbound_verify/sites/anneng.py` | 同上 |
| `inbound_verify/sites/shunxin.py` | `download/impl``force`方式1点运单列表后读交接单号去重 + 退回 + 空兜底 |
| `dashboard/app/page.tsx` | 加 `forceRedownload` state + checkbox`trigger/triggerPrimary` 透传 force |
> 4 站的 actual实到`download` 入口也统一加 `force=False` 形参(接收但不用,仅让 `_web_handler` 的统一调用成立actual impl 不改。
## 5. 任务分解
### Task 1schema.sql 加索引
**Files:** Modify `schema.sql``idx_expected_site_date` 之后)
```sql
CREATE INDEX IF NOT EXISTS idx_expected_handover ON expected_record (site, handover_no);
```
**验证:** `.venv/Scripts/python.exe -m inbound_verify.store init`(幂等)。
---
### Task 2store.py 新增查已落库交接单号集合
**Files:** Modify `inbound_verify/store.py``ingest_task` 之后)
**Produces:** `get_existing_handover_nos(site: str) -> set[str]`
```python
def get_existing_handover_nos(site):
"""查该站点已落库的交接单号集合expected_record.handover_no
"提交导出任务前"去重:已落库的不再重复提交导出。
PG 不可用cpolar 抖动等)时返回空集 + 告警,调用方按"未确认存在"处理
继续提交导出UPSERT 兜底,绝不因去重查询失败而漏数据)。"""
try:
with _connect(_load_pg_config()["dbname"]) as conn:
with conn.cursor() as cur:
cur.execute(
"SELECT handover_no FROM expected_record "
"WHERE site=%s AND handover_no IS NOT NULL AND handover_no <> ''",
(site,),
)
return {str(r[0]).strip() for r in cur.fetchall()}
except Exception as e:
print(f">> [去重] 查询已落库交接单号失败({site}),本次不去重: {e}")
return set()
```
**验证:** `.venv/Scripts/python.exe -c "from inbound_verify import store; print(len(store.get_existing_handover_nos('中通')))"` 不抛异常即过。
---
### Task 3force 开关后端骨架server + runtime
`force``task_spec` 一路透传到各站 `download(page, force)``with_retry``flow` 是零参 lambdaforce 经闭包捕获,**with_retry 不动**。
**(a) `server.py` TaskRequest + create_task**
```python
class TaskRequest(BaseModel):
site: str
kind: str
force: bool = False # 新增:强制重下(忽略已落库去重),默认关
```
```python
# create_task 内
task_queue.put((task_id, {"site": req.site, "kind": req.kind, "force": req.force}))
```
> `_enqueue_fetch`L133 周期投递)**保持不变**(不带 force → 默认 False✅。
**(b) `runtime.py` dispatch_task 透传 force**
```python
def dispatch_task(ctx, task_spec):
site = task_spec.get("site")
kind = task_spec.get("kind")
force = bool(task_spec.get("force", False)) # 新增
...
try:
ret = handler(ctx, force) # 改:原 handler(ctx)
```
**(c) `runtime.py` 所有 handler 加 force 形参:**
`_web_handler`
```python
def handler(ctx, force=False):
pg = ctx.pages_map[site]
if isinstance(pg, list): # 顺心双账号
return download_func(pg, foreground=ctx.foreground, force=force)
if ctx.foreground:
pg.bring_to_front()
return download_func(pg, force=force)
```
`_site_undelivered_handler`
```python
def handler(ctx, force=False):
exp_ok = TASK_HANDLERS[(site, "expected")](ctx, force) is not False
act_ok = (TASK_HANDLERS[(site, "actual")](ctx, force) is not False) if exp_ok else False
...
```
安能 expected/actual 与 compare签名兼容即可
```python
("安能", "expected"): lambda ctx, force=False: anneng.anneng_expected_download(force=force),
("安能", "actual"): lambda ctx, force=False: anneng.anneng_actual_download(force=force),
("__compare__", "compare"): lambda ctx, force=False: (compare.main() or True),
```
**验证:** `py_compile` 通过;服务重启后 `POST /tasks {site,kind,force:true}` 不报 TypeError此时各站 `download` 的 force 形参由 Task 4-7 补齐,连续实施)。
---
### Task 4中通 zto.pydownload/impl 加 force + 去重)
**Files:** Modify `inbound_verify/sites/zto.py`
**(a) 入口与 impl 加 force闭包透传with_retry 不动):**
```python
def zto_expected_download(page, force=False):
return with_retry(
"中通", "应到",
lambda: zto_expected_download_impl(page, force=force),
lambda: zto_reset(page),
)
def zto_expected_download_impl(page, force=False):
...
# zto_actual_download / zto_actual_download_impl 同样加 force=False 形参actual 不用 force仅兼容
```
**(b) 循环前加载已落库集合**L241 print 之后、L243 `for` 之前):
```python
# 【去重】加载本站已落库交接单号force=True 或查询失败时 existing=空集(不去重)
if force:
existing = set()
print(">> [去重] 强制重下,跳过去重。")
else:
try:
from inbound_verify import store
existing = store.get_existing_handover_nos("中通")
except Exception as _e:
existing = set()
print(f">> [去重] 加载失败,本次不去重: {_e}")
```
**(c) 循环内命中跳过**L252 print 之后、L254 `row.dblclick()` 之前):
```python
print(f" -> 当前交接单号:{handover_no}")
if handover_no in existing:
print(f" ⏭️ 交接单号 {handover_no} 已落库,跳过提交导出。")
continue
row.dblclick()
```
**(d) 空列表兜底**L305 循环后、进入 `_zto_poll_and_download_tasks` 之前):
```python
if not export_times:
print(">> 本次无新交接单需导出(全部已落库或无数据),结束。")
return
```
---
### Task 5韵达 yunda.py同构
**Files:** Modify `inbound_verify/sites/yunda.py`
(a) `yunda_expected_download(page, force=False)` + impl 加 forceactual 同理加形参);(b) 循环前加载 existing同 Task 4b站点"韵达"
**(c) 循环内命中跳过**L291 `raw_no = ...` 之后、L293 `# 跳过已绑定的交接单` 之前):
```python
raw_no = current_row.locator("td").nth(1).inner_text().strip()
if raw_no in existing:
print(f" ⏭️ 交接单号 {raw_no} 已落库,跳过提交导出。")
continue
# 跳过已绑定的交接单
bind_status = current_row.locator("td").nth(2).inner_text().strip()
```
(d) 空列表兜底 **已存在**L389-392无需新增。
---
### Task 6安能 anneng.pyCDP
**Files:** Modify `inbound_verify/sites/anneng.py`
(a) `anneng_expected_download(force=False)` + impl 加 forceactual 同理);
**(b) 主流程加载 existing**`anneng_expected_download_impl` 内、L907 收集 `target_ids` 前,与 `export_times = []`L882并列同 Task 4b站点"安能"
**(c) 主循环命中跳过**L919 `for` 内、L920 print 之后、L921 `activate_tab` 之前):
```python
for i, ewbs_no in enumerate(target_ids, start=1):
print(f" ⏳ [{i}/{len(target_ids)}] 交接单号 {ewbs_no}")
if ewbs_no in existing:
print(f" ⏭️ 交接单号 {ewbs_no} 已落库,跳过。")
continue
activate_tab(tab_cdp, "交接单信息")
...
```
**(d) 空列表兜底**(主循环后、进入 `poll_and_download_tasks` 之前):
```python
if not export_times:
print(">> 本次无新交接单需导出(全部已落库或无数据),结束。")
return
```
---
### Task 7顺心 shunxin.py方式1 + 实勘)
顺心流程L315 点"运单列表" → L316 等"运单查询"label**运单列表界面** → L319 点"导出"。方式1 在 L316 之后、L319 之前读交接单号。
**Step 1实勘不改业务逻辑** 确认运单列表界面里**交接单号的 DOM 选择器**(哪个元素/列)。实勘方式二选一(待用户同意):
- 方式 A临时在 L316 后加调试打印dump 运单列表界面关键 DOM 文本),跑一次顺心应到,从日志定位选择器,再删调试代码。
- 方式 Bdebug 模式(`config.yaml` debug.target_site=顺心CDP 9223单独挂载用 Playwright CLI 观察。
**Step 2注入选择器 `<HANDOVER_SELECTOR>` 确认后替换):**
(a) 入口与 impl 加 force
```python
def shunxin_expected_download(pages, foreground=True, force=False):
return with_retry(
"顺心", "应到",
lambda: shunxin_expected_download_impl(pages, foreground=foreground, force=force),
lambda: shunxin_reset(pages),
)
def shunxin_expected_download_impl(pages, foreground=True, force=False):
...
# shunxin_actual_download / impl 同样加 force=False 形参actual 不用)
```
(b) 循环前加载 existing同 Task 4b站点"顺心";两账号共享同一 `existing`)。
**(c) 循环内:点运单列表 → 读交接单号 → 命中则退回跳过**L315-316 之后、L319 点导出之前):
```python
for i in range(count):
print(f" ⏳ 正在处理第 {i+1}/{count} 个班次...")
waybill_btns.nth(i).click()
page.locator("label[title='运单查询']").wait_for(state="visible")
# 【方式1】运单列表界面已加载读交接单号 → 已落库则退回列表跳过
handover_no = page.locator("<HANDOVER_SELECTOR>").first.inner_text().strip()
if handover_no in existing:
print(f" ⏭️ 交接单号 {handover_no} 已落库,跳过提交导出。")
page.get_by_role("tab", name="车辆点到").click() # 退回列表(复用 L336
page.wait_for_timeout(500)
continue
# 4. 执行导出流程
page.get_by_role("button", name="export 导出").click()
...
```
(d) 空列表兜底L339 循环后、进入导出任务管理页轮询之前):
```python
if not export_times:
print(">> 本次无新班次需导出(全部已落库或无数据),结束。")
return
```
---
### Task 8前端 force 开关dashboard
**Files:** Modify `dashboard/app/page.tsx`BFF `app/api/tasks/route.ts` 是 generic 透传,**不用改**
**(a) 加 stateL21 附近):**
```tsx
const [forceRedownload, setForceRedownload] = useState(false);
```
**(b) `trigger` 加 force 形参并写入 bodyL27-50**
```tsx
const trigger = useCallback(
async (site: string, kind: string, label: string, force?: boolean) => {
...
body: JSON.stringify({ site, kind, force: !!force }),
...
},
[refreshTasks],
);
```
**(c) `triggerPrimary` 透传L58-63**
```tsx
const triggerPrimary = useCallback(
async (cfg: SiteConfig) => {
await trigger(cfg.name, cfg.primaryKind, `${cfg.name}·获取未到`, forceRedownload);
},
[trigger, forceRedownload],
);
```
**(d) UI在配置区/顶部加 checkbox默认不勾**
```tsx
<label className="inline-flex items-center gap-1 text-xs text-amber-700">
<input
type="checkbox"
checked={forceRedownload}
onChange={(e) => setForceRedownload(e.target.checked)}
/>
</label>
```
> 仅"获取未到"主按钮透传 force周期抓取不经前端、恒不 force。
**验证:** 前端勾选 → 触发 → 网络面板看到 POST `/api/tasks` body 含 `force:true`;后端日志 `强制重下,跳过去重`
## 6. 关键注意事项(陷阱)
- **跳过的交接单号绝不 `export_times.append`**:否则 `len(export_times)` > 实际提交数 → 下载数校验失败。所有 `continue` 都在 append 之前。
- **全部跳过时必须 `return`**`export_times` 为空时不进导出任务管理页轮询。
- **顺心方式1跳过要退回**:点进运单列表后命中已落库,需点"车辆点到"tab 退回再 `continue`(复用 L336
- **PG 降级只防漏不防重**:查询失败 = 空集 = 当作未存在 = 继续提交UPSERT 兜底。
- **不动 `export_times` 时间容差匹配**。
- actual/百世 `download` 只加 `force` 形参兼容impl 不加去重。
## 7. 验证(无 pytest
1. **单站联调**`config.yaml` debug 单站):首次新单正常下载+入库;再触发同范围 → 已入库的全部 `⏭️ 跳过``export_times` 空,直接 return。
2. **force 开关**:勾选"强制重下" → 已落库的也重新提交导出(日志 `强制重下,跳过去重`)。
3. **DB 核对**`SELECT site, handover_no, COUNT(*) FROM expected_record GROUP BY site, handover_no` 无翻倍。
4. **cpolar 降级**:断 PG → `get_existing_handover_nos` 返回空集 + 告警,流程仍正常下载(不漏)。
5. **Black + py_compile**;前端 `npm run build` 或 dev 热更无类型错。
## 8. 实施顺序与依赖
Task 1 → 2 → 3骨架此时各站 download 的 force 形参在 4-7 补)→ 4/5/6/7各站连续做完让链路自洽→ 8前端。顺心 Task 7 的 Step1 实勘需在运行的服务上操作,实施时与用户协调时机。

View File

@@ -0,0 +1,131 @@
# 指定日期下载接口(开发者)— 设计文档
> 日期2026-07-29
> 定位:面向开发者的 HTTP 接口,**不进前端**。提供"指定一个具体日期,下载该日应到 / 实到数据"的能力,用于补下历史数据。
> 前置:中通跨月导航已实现并验证(见 `feat(zto): cross-month calendar navigation`)。
## 一、背景与目标
现状:下载日期由各站 **offset 偏移**0=今天1=昨天…,存 `state.db`,上限 `MAX_DATE_OFFSET=30`)决定,周期调度与手动触发都用 offset。无法指定一个具体日期。
目标:新增"指定日期"入口(开发者用),传一个 `YYYY-MM-DD` 日期,下载该日应到 / 实到数据。不替换 offset 机制,与之并存:传 date 用 date不传走 offset。
## 二、范围
| 站点 | 支持指定日期 | 说明 |
| --- | --- | --- |
| 顺心 / 中通 / 韵达 / 安能 | ✅ | 应到、实到均支持 |
| 百世 | ❌ | 固定下载当天,传 date 返回 400 |
## 三、接口契约
复用 `POST /tasks`body 新增可选字段 `date`(与 `force` 并列):
```json
{ "site": "中通", "kind": "expected", "date": "2026-06-14" }
```
- `date: Optional[str] = None`,格式 `YYYY-MM-DD`
- **优先级**:传 `date` 则本次用 date不传则走站点 offset 配置(默认行为完全不变)。
- `date``force` 可共存(指定日期 + 强制重下)。
### 合法性校验(仅在传了 date 时执行,失败返回 400
1. **格式**`datetime.strptime(date, "%Y-%m-%d")` 解析成功,否则 400。
2. **范围**`今天 - 90 天 ≤ date ≤ 今天`
- `date > 今天` → 400未来日期日历未来格子 `invalid` 物理上点不动,且不应下未来数据)。
- `date < 今天 - 90 天` → 400回溯上限 90 天)。
3. **百世**site=百世 且传 date → 400固定当天
> 合法性校验落在 `POST /tasks``server.py` `create_task`),入队前拦截,非法请求不产生任务。
## 四、透传链路(与现有 `force` 完全对称)
```
POST /tasks {site, kind, force, date}
→ task_queue.put((tid, {site, kind, force, date}))
→ dispatch_task(ctx, task_spec) # 读 task_spec["date"]
→ handler(ctx, force, date) # _web_handler / 安能 lambda / _site_undelivered_handler
→ impl(page, force, date) # 各站 download_impl
```
- `_web_handler``handler(ctx, force=False, date=None)`,透传 `download_func(pg, force, date)`;顺心双账号透传 `(pages, foreground, force, date)`
- `_site_undelivered_handler`(未到):连下 expected + actual**两个子任务共用同一个 date**。
- 安能 lambda`(ctx, force=False, date=None) → anneng_xxx_download(force=force, date=date)`
- **周期调度** `_enqueue_fetch` 投递的 task_spec 只有 `{site, kind}`(不带 date→ 恒走 offset**无需改动**。
## 五、各站 impl 改造(核心)
统一模式:**`target = parse(date) if date else (today offset)`**。
### 中通zto—— 复用跨月算法
把 date 折算成 effective offset复用现有 `target_time = today_time offset*86400000` 与跨月翻页(`_zto_flip_to_target_month`),零额外 UI 逻辑:
```python
def zto_expected_download_impl(page, force=False, date=None):
...
offset = state_store.get_offset("中通")
if date:
target_date = datetime.strptime(date, "%Y-%m-%d").date()
offset = (datetime.now().date() - target_date).days
# 后续 today_time / target_time / 跨月翻页 逻辑完全不变
```
`zto_actual_download_impl` 同理(用 `("中通","actual")` offset。expected / actual 两个 impl 都加 `date=None` 形参,`zto_expected_download` / `zto_actual_download` 公开入口同步加形参并透传。
### 韵达 / 顺心 / 安能 —— date 直接当 target
这三站 offset→日期是 `target = today timedelta(days=offset)` 后填**字符串**到日期控件(非日历格子),指定日期只需替换 target 来源:
```python
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
```
后接的"填起始/截止日期字符串"逻辑完全不变。各站 expected / actual 入口与 impl 都加 `date=None` 形参。
- 韵达:`yunda_expected_download(_impl)` / `yunda_actual_download(_impl)`
- 顺心:`shunxin_expected_download(_impl)` / `shunxin_actual_download(_impl)`(双账号入口透传 date 到各账号 impl
- 安能:`anneng_expected_download` / `anneng_actual_download`
### 百世baishi—— 签名兼容
`baishi_download_undelivered_data(page, date=None)``date=None` 形参(**忽略**),仅为对齐 `_web_handler` 的统一透传签名;百世任务实际不会带 dateserver 已拦截)。
## 六、业务日期快照
`_record_business_date(site, kind, date=None)`:有 date 则业务日期 = date否则维持现状 `today offset``dispatch_task``task_spec["date"]` 透传进去,保证状态盘 / 报告显示的"是哪天的数据"准确(不被 offset 算错)。
## 七、改动文件清单
| 文件 | 改动 |
| --- | --- |
| `inbound_verify/cli/server.py` | `TaskRequest.date` + `create_task` 合法性校验 + task_spec 透传 date |
| `inbound_verify/runtime.py` | `_web_handler` / `_site_undelivered_handler` / 安能 lambda 透传 date`dispatch_task` 读 date 透传给 handler 与 `_record_business_date``_record_business_date` 加 date |
| `inbound_verify/sites/zto.py` | expected/actual 入口+impl 加 `date`date→effective offset 复用跨月 |
| `inbound_verify/sites/yunda.py` | expected/actual 入口+impl 加 `date`date→target |
| `inbound_verify/sites/shunxin.py` | 同上(双账号透传 date |
| `inbound_verify/sites/anneng.py` | expected/actual 加 `date`date→target |
| `inbound_verify/sites/baishi.py` | 加 `date=None` 形参兼容(忽略) |
## 八、验证计划
1. **接口校验**curl/python urllib
- 合法 date过去某日→ 202任务成功。
- 未来日期 / 超 90 天 / 格式错 → 400。
- 百世 + date → 400。
- 不传 date → 走 offset行为不变
2. **各站实测**(指定一个过去日期触发任务):
- 中通:跨月日期(已知 OK复用已验证的跨月导航
- 顺心 / 韵达 / 安能:实测其日期控件是否接受任意过去日期字符串;若控件是日历选择器需翻月,则按中通同法扩展(本轮发现则记录、必要时追加改动)。
3. **业务日期快照**:下载后 `GET /status``*_business_date` == 指定 date。
4. 改完跑 Black + `py_compile`
## 九、非目标YAGNI
- 前端 UIcheckbox / 日期选择器)——开发者接口,不进前端。
- 周期调度指定日期——周期恒走 offset。
- 批量日期 / 日期范围下载——单次单日。

View File

@@ -0,0 +1,682 @@
# 指定日期下载接口开发者Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:**`POST /tasks` 加可选 `date`YYYY-MM-DD支持顺心/中通/韵达/安能指定一个过去日期下载应到/实到数据;不传 date 时行为不变(走 offset
**Architecture:** 复用现有 `force` 透传链路加一个 `date` 字段:`TaskRequest.date``task_spec["date"]``dispatch_task``handler(ctx, force, date)` → 各站 `impl(page, force, date)`。各站把"算 target 日期"的来源从 `today - offset` 改为"有 date 用 date否则 today - offset"。中通把 date 折算成 effective offset 以复用已验证的跨月翻页。
**Tech Stack:** Python 3.10、FastAPI、Playwright、SQLitestate_store、PostgreSQLstore可选入库
## Global Constraints
- **环境**:所有 Python 用 `D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe`项目虚拟环境无需激活。PYTHONPATH 含 `InboundVerify` 根。
- **格式化/自检(每改一个 .py 必做)**`python -m black <file>`target py310+ `python -m py_compile <file>`。Black 若提示 "Python 3.10 cannot parse code formatted for 3.15",加 `--target-version py310`"left unchanged" 即合格。
- **无 pytest**:本项目无单元测试框架(见 `InboundVerify/CLAUDE.md`)。每个任务的"验证"= Black + py_compile端到端API 校验、各站实测)集中在 Task 8需重启服务加载新代码
- **submodule 工作流**:改动在 `InboundVerify` submodule`dev` 分支)。每个 Task 末尾在 submodule 内 `git add <file> && git commit`。**push 到 origin/dev + 父仓库 bump** 统一在 Task 8项目约定不自动 push等用户确认但 plan 内 commit 步骤照写)。
- **顺序安全**Task 1-5站点只给 impl 加 `date=None` 形参 + date 逻辑,**date 默认 None 时走原 offset 路径,向后兼容**Task 6runtime才把 date 从 task_spec 透传进 implTask 7server才允许 date 入队。任一 Task 完成后系统均可正常运行。
- **合法性边界**(来自 spec`date` 传了才校验——格式 `YYYY-MM-DD``今天-90 ≤ date ≤ 今天`、百世不支持 date。非法返回 HTTP 400。
---
## File Structure
| 文件 | 责任 | 本计划改动 |
| --- | --- | --- |
| `inbound_verify/sites/baishi.py` | 百世下载(固定当天) | 入口加 `date=None` 形参(忽略) |
| `inbound_verify/sites/zto.py` | 中通下载(日历格子,跨月) | 入口+impl 加 `date`date→effective offset 复用跨月 |
| `inbound_verify/sites/yunda.py` | 韵达下载(日期字符串) | expected/actual 入口+impl 加 `date`date→target |
| `inbound_verify/sites/shunxin.py` | 顺心下载(双账号,日期字符串) | expected/actual 入口+impl 加 `date`date→target |
| `inbound_verify/sites/anneng.py` | 安能下载CDP日期字符串 | expected/actual 入口+impl 加 `date`date→target |
| `inbound_verify/runtime.py` | 任务派发/心跳共享核心 | handler/dispatch 透传 date`_record_business_date` 用 date |
| `inbound_verify/cli/server.py` | FastAPI 服务 | `TaskRequest.date` + 合法性校验 + task_spec 透传 |
---
## Task 1: baishi.py — 入口加 date=None 形参(兼容)
**Files:**
- Modify: `inbound_verify/sites/baishi.py:140``baishi_download_undelivered_data`
**Interfaces:**
- Produces: `baishi_download_undelivered_data(page, force=False, date=None)` —— 后续 Task 6 的 `_web_handler` 会以 `download_func(pg, force=force, date=date)` 调用它,必须接受 `date` kwarg百世忽略
- [ ] **Step 1: 改签名**
`def baishi_download_undelivered_data(page, force=False):` 改为:
```python
def baishi_download_undelivered_data(page, force=False, date=None):
"""百世应到未到数据下载固定当天date 形参仅为对齐统一透传签名,忽略)。"""
```
(函数体不动;`date` 不使用。)
- [ ] **Step 2: Black + py_compile**
```bash
PY="D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe"
F="D:/projects/LogisticsHubIPA/InboundVerify/inbound_verify/sites/baishi.py"
"$PY" -m black "$F" && "$PY" -m py_compile "$F" && echo OK
```
Expected: black "left unchanged" 或 reformat 后通过OK。
- [ ] **Step 3: Commit**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" add inbound_verify/sites/baishi.py
git -C "D:/projects/LogisticsHubIPA/InboundVerify" commit -m "feat(baishi): accept date kwarg (ignored) for unified dispatch signature" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 2: zto.py — date 折算成 effective offset复用跨月
**Files:**
- Modify: `inbound_verify/sites/zto.py``zto_expected_download``zto_expected_download_impl``zto_actual_download``zto_actual_download_impl`
**Interfaces:**
- Produces: `zto_expected_download(page, force=False, date=None)` / `zto_actual_download(page, force=False, date=None)`impl 同签名。Task 6 的 `_web_handler``download_func(pg, force=force, date=date)` 调用。
- [ ] **Step 1: expected 入口加 date 并透传**
`def zto_expected_download(page, force=False):` 及其 `with_retry` 改为:
```python
def zto_expected_download(page, force=False, date=None):
"""中通:应到货物数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry(
"中通",
"应到",
lambda: zto_expected_download_impl(page, force=force, date=date),
lambda: zto_reset(page),
)
```
- [ ] **Step 2: expected impl 加 datedate→effective offset**
`def zto_expected_download_impl(page, force=False):` 改签名加 `date=None`。其内"读取服务端日期偏移"段(`offset = state_store.get_offset("中通")` 与随后的 `print(...偏移...)`)改为:
```python
# 读取服务端日期偏移0=今天1=昨天…),单日:起止同日
offset = state_store.get_offset("中通")
if date:
# 指定日期:折算成相对今天的有效偏移,复用下方 target_time 计算与跨月翻月
target_date = datetime.strptime(date, "%Y-%m-%d").date()
offset = (datetime.now().date() - target_date).days
print(f">> 正在设定查询日期: 指定日期 {date}(折算偏移 {offset}...")
else:
print(f">> 正在设定查询日期: 偏移 {offset}0=今天)...")
```
(其后的 `target_time = today_time - offset * 86400000` 与跨月翻月逻辑**不动**——date 经折算后走同一条路径。)
- [ ] **Step 3: actual 入口加 date 并透传**
`def zto_actual_download(page, force=False):` 及其 `with_retry` 改为:
```python
def zto_actual_download(page, force=False, date=None):
"""中通:实到货物数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry(
"中通",
"实到",
lambda: zto_actual_download_impl(page, date=date),
lambda: zto_reset(page),
)
```
> 注:`zto_actual_download_impl` 现签名 `(page)`(无 forceactual 不去重Step 4 给它加 `date`。
- [ ] **Step 4: actual impl 加 datedate→effective offset**
`def zto_actual_download_impl(page):` 改为 `def zto_actual_download_impl(page, date=None):`。其内"# 2. 读取服务端日期偏移"段(`offset = state_store.get_offset("中通", "actual")` 与随后的 `print`)改为:
```python
# 2. 读取服务端日期偏移0=今天1=昨天…),单日:起止同日
offset = state_store.get_offset("中通", "actual")
if date:
target_date = datetime.strptime(date, "%Y-%m-%d").date()
offset = (datetime.now().date() - target_date).days
print(f">> 正在设定查询日期: 指定日期 {date}(折算偏移 {offset}...")
else:
print(f">> 正在设定查询日期: 偏移 {offset}0=今天)...")
```
(其后 `target_time = today_time - offset * 86400000` 与跨月翻月不动。)
- [ ] **Step 5: Black + py_compile**
```bash
PY="D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe"
F="D:/projects/LogisticsHubIPA/InboundVerify/inbound_verify/sites/zto.py"
"$PY" -m black "$F" && "$PY" -m py_compile "$F" && echo OK
```
- [ ] **Step 6: Commit**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" add inbound_verify/sites/zto.py
git -C "D:/projects/LogisticsHubIPA/InboundVerify" commit -m "feat(zto): support date arg via effective-offset (reuses cross-month nav)" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 3: yunda.py — date→targetexpected + actual
**Files:**
- Modify: `inbound_verify/sites/yunda.py``yunda_expected_download(_impl)``yunda_actual_download(_impl)`
**Interfaces:**
- Produces: `yunda_expected_download(page, force=False, date=None)` / `yunda_actual_download(page, force=False, date=None)`impl 同加 `date=None`
- [ ] **Step 1: expected 入口透传 date**
```python
def yunda_expected_download(page, force=False, date=None):
...
return with_retry(
"韵达",
"应到",
lambda: yunda_expected_download_impl(page, force=force, date=date),
lambda: yunda_reset(page),
)
```
- [ ] **Step 2: expected impl 加 date + date→target**
`def yunda_expected_download_impl(page, force=False):``def yunda_expected_download_impl(page, force=False, date=None):`。其内日期段(`offset = state_store.get_offset("韵达")` 起几行)改为:
```python
offset = state_store.get_offset("韵达")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_ymd = f"{target.year}-{target.month}-{target.day}"
start_date_ymd = target_ymd
today_ymd = target_ymd
print(f">> 设置查询日期: [{target_ymd}]{'指定 ' + date if date else f'偏移 {offset}0=今天'}")
```
- [ ] **Step 3: actual 入口透传 date**
```python
def yunda_actual_download(page, force=False, date=None):
...
return with_retry(
"韵达",
"实到",
lambda: yunda_actual_download_impl(page, date=date),
lambda: yunda_reset(page),
)
```
- [ ] **Step 4: actual impl 加 date + date→target**
`def yunda_actual_download_impl(page):``def yunda_actual_download_impl(page, date=None):`。其内日期段改为:
```python
offset = state_store.get_offset("韵达", "actual")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
start_date = target # 单日范围:起止同日
today = target # 让下方"截止时间"选择器也指向 target
print(
f">> 设置实到查询日期: [{target.year}-{target.month}-{target.day}]"
f"{'指定 ' + date if date else f'偏移 {offset}0=今天'}"
)
```
- [ ] **Step 5: Black + py_compile**(同 Task 2 命令,文件换 yunda.py
- [ ] **Step 6: Commit**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" add inbound_verify/sites/yunda.py
git -C "D:/projects/LogisticsHubIPA/InboundVerify" commit -m "feat(yunda): support date arg (date takes precedence over offset)" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 4: shunxin.py — date→target双账号透传
**Files:**
- Modify: `inbound_verify/sites/shunxin.py``shunxin_expected_download(_impl)``shunxin_actual_download(_impl)`
**Interfaces:**
- Produces: `shunxin_expected_download(pages, foreground=True, force=False, date=None)` / `shunxin_actual_download(pages, foreground=True, force=False, date=None)`。Task 6 的 `_web_handler``download_func(pg, foreground=ctx.foreground, force=force, date=date)` 调用pg 是 page 列表)。
- [ ] **Step 1: expected 入口加 date 并向 impl 透传**
`def shunxin_expected_download(pages, foreground=True, force=False):` → 加 `date=None`。在函数内调用 `shunxin_expected_download_impl(page, out_tag=..., force=force)` 的位置,加上 `date=date`
```python
def shunxin_expected_download(pages, foreground=True, force=False, date=None):
...
# 对每个账号调用 impl 时透传 date
... shunxin_expected_download_impl(page, out_tag=归属地, force=force, date=date) ...
```
> 执行者:用 Grep 定位 `shunxin_expected_download_impl(page,` 的调用处(在 `shunxin_expected_download` 函数体内,对每个账号调用一次),给每处加 `date=date`。签名行只加 `date=None`其余函数体归属地读取、去重校验、merge不动。
- [ ] **Step 2: expected impl 加 date + date→target**
`def shunxin_expected_download_impl(page, out_tag="", force=False):` → 加 `date=None`。其内日期段(`offset = state_store.get_offset("顺心")` 起几行)改为:
```python
offset = state_store.get_offset("顺心")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_str = target.strftime("%Y-%m-%d")
start_date_str = target_str
today_str = target_str
print(f">> 正在设置查询日期: [{target_str}]{'指定 ' + date if date else f'偏移 {offset}0=今天'}...")
```
- [ ] **Step 3: actual 入口加 date 并透传**
`def shunxin_actual_download(pages, foreground=True, force=False):` → 加 `date=None`;调用 `shunxin_actual_download_impl(page, out_tag=...)` 处加 `date=date`
- [ ] **Step 4: actual impl 加 date + date→target**
`def shunxin_actual_download_impl(page, out_tag=""):``def shunxin_actual_download_impl(page, out_tag="", date=None):`。其内日期段(`offset = state_store.get_offset("顺心", "actual")` 起几行)改为:
```python
offset = state_store.get_offset("顺心", "actual")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_str = target.strftime("%Y-%m-%d")
start_date_str = target_str
today_str = target_str
print(f">> 正在设置查询日期: [{target_str}]{'指定 ' + date if date else f'偏移 {offset}0=今天'}...")
```
- [ ] **Step 5: Black + py_compile**(文件 shunxin.py
- [ ] **Step 6: Commit**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" add inbound_verify/sites/shunxin.py
git -C "D:/projects/LogisticsHubIPA/InboundVerify" commit -m "feat(shunxin): support date arg, propagate to both accounts" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 5: anneng.py — date→targetCDPexpected + actual
**Files:**
- Modify: `inbound_verify/sites/anneng.py``anneng_expected_download(_impl)``anneng_actual_download(_impl)`
**Interfaces:**
- Produces: `anneng_expected_download(force=False, date=None)` / `anneng_actual_download(force=False, date=None)`。Task 6 的 TASK_HANDLERS 安能项以 `lambda ctx, force=False, date=None: anneng.anneng_expected_download(force=force, date=date)` 调用。
- [ ] **Step 1: expected 入口 + impl 加 datedate→target**
```python
def anneng_expected_download(force=False, date=None):
return with_retry(
"安能", "应到", lambda: anneng_expected_download_impl(force=force, date=date), anneng_reset
)
def anneng_expected_download_impl(force=False, date=None):
```
expected impl 内日期段(`offset = state_store.get_offset("安能")` 起几行)改为:
```python
offset = state_store.get_offset("安能")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_str = f"{target.year}-{target.month:02d}-{target.day:02d}"
start_str = target_str
today_str = target_str
print(f">> 查询日期: [{target_str}]{'指定 ' + date if date else f'偏移 {offset}0=今天'}")
```
- [ ] **Step 2: actual 入口 + impl 加 datedate→target**
```python
def anneng_actual_download(force=False, date=None):
return with_retry("安能", "实到", lambda: anneng_actual_download_impl(date=date), anneng_reset)
def anneng_actual_download_impl(date=None):
```
actual impl 内日期段(`offset = state_store.get_offset("安能", "actual")` 起几行)改为:
```python
offset = state_store.get_offset("安能", "actual")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
start_str = f"{target.year}\{target.month:02d}/{target.day:02d} 00:00:00"
end_str = f"{target.year}\{target.month:02d}/{target.day:02d} 23:59:59"
print(f">> 扫描日期: [{start_str}{end_str}]{'指定 ' + date if date else f'偏移 {offset}0=今天'}")
```
> actual 的 `start_str/end_str` 沿用现有 `{year}\{month}/{day}` 格式(含反斜杠,站点如此),只把 target 来源改成 date。
- [ ] **Step 3: Black + py_compile**(文件 anneng.py
- [ ] **Step 4: Commit**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" add inbound_verify/sites/anneng.py
git -C "D:/projects/LogisticsHubIPA/InboundVerify" commit -m "feat(anneng): support date arg (date takes precedence over offset)" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 6: runtime.py — 全链路透传 date + _record_business_date 用 date
**Files:**
- Modify: `inbound_verify/runtime.py``_web_handler``_site_undelivered_handler``TASK_HANDLERS` 安能项、`dispatch_task``_record_business_date`
**Interfaces:**
- Consumes: Task 1-5 产出的各站 `*(..., date=None)` 签名。
- Produces: `dispatch_task``task_spec["date"]` 读 date 透传给 handler 与 `_record_business_date`handler 签名 `(ctx, force=False, date=None)`
- [ ] **Step 1: _web_handler 透传 date**
```python
def _web_handler(site, download_func):
def handler(ctx, force=False, date=None):
pg = ctx.pages_map[site]
if isinstance(pg, list):
# 顺心双账号:置顶与否交给 shunxin_download 在逐账号循环里按 foreground 决定
return download_func(pg, foreground=ctx.foreground, force=force, date=date)
if ctx.foreground:
pg.bring_to_front()
return download_func(pg, force=force, date=date)
return handler
```
- [ ] **Step 2: _site_undelivered_handler 透传 date连下 expected+actual 共用同一 date**
```python
def _site_undelivered_handler(site):
def handler(ctx, force=False, date=None):
exp_ok = TASK_HANDLERS[(site, "expected")](ctx, force, date) is not False
act_ok = (
(TASK_HANDLERS[(site, "actual")](ctx, force, date) is not False)
if exp_ok
else False
)
if exp_ok and act_ok:
return compare.write_site_file(site)
stale = os.path.join(DOWNLOAD_DIR, SITE_UNDELIVERED_FILE.format(name=site))
if os.path.exists(stale):
os.remove(stale)
return False
return handler
```
- [ ] **Step 3: TASK_HANDLERS 安能项透传 date**
```python
("安能", "expected"): lambda ctx, force=False, date=None: anneng.anneng_expected_download(
force=force, date=date
),
("安能", "actual"): lambda ctx, force=False, date=None: anneng.anneng_actual_download(
force=force, date=date
),
```
`__compare__` 项与百世/网页项不动——百世经 `_web_handler` 已透传 datebaishi 忽略。)
- [ ] **Step 4: dispatch_task 透传 date**
`dispatch_task` 内,把 `ret = handler(ctx, bool(task_spec.get("force", False)))` 改为:
```python
ret = handler(
ctx, bool(task_spec.get("force", False)), task_spec.get("date")
)
if ret is False:
return (state_store.TASK_FAILED, "任务执行失败(重试耗尽)")
_record_business_date(site, kind, task_spec.get("date"))
_persist_to_db(site, kind)
return (state_store.TASK_SUCCESS, None)
```
- [ ] **Step 5: _record_business_date 接受 date**
```python
def _record_business_date(site, kind, date=None):
"""下载成功后,把本次数据的业务日期快照写进状态库(供前端/报告显示「是哪天的数据」)。
有 date 用 date否则 = 下载当天 该数据对应的日期偏移。__compare__ 无数据概念,跳过。"""
if site == "__compare__":
return
today = datetime.now().date()
now = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
def _write(k, biz_or_off):
# biz_or_off: int=offsettodayoffsetstr=已确定的业务日期date
biz = (
(today - timedelta(days=biz_or_off)).strftime("%Y-%m-%d")
if isinstance(biz_or_off, int)
else biz_or_off
)
try:
state_store.set_data_state(
site, k, ready=True, generated_at=now, business_date=biz
)
except Exception as e:
print(f">> [状态] 写业务日期失败 {site}/{k}: {e}")
def off(kind_key):
return state_store.get_offset(site, kind_key)
if kind == "expected":
_write("expected", date if date else off("expected"))
elif kind == "actual":
_write("actual", date if date else off("actual"))
elif site == "百世":
_write("undelivered", 0)
else: # 4 站 undelivered连带补写 expected/actual/undelivered 三列
_write("expected", date if date else off("expected"))
_write("actual", date if date else off("actual"))
_write("undelivered", date if date else off("expected"))
```
- [ ] **Step 6: Black + py_compile**
```bash
PY="D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe"
F="D:/projects/LogisticsHubIPA/InboundVerify/inbound_verify/runtime.py"
"$PY" -m black "$F" && "$PY" -m py_compile "$F" && echo OK
```
- [ ] **Step 7: 冒烟date=None 行为不变)**
服务仍跑旧代码,但 runtime 模块可独立 import 校验:
```bash
PYTHONPATH="D:/projects/LogisticsHubIPA/InboundVerify" "$PY" -c "import inbound_verify.runtime as r; import inspect; print('handler', inspect.signature(r._web_handler('中通', lambda *a, **k: None))); print('rec', inspect.signature(r._record_business_date))"
```
Expected: handler 含 `(ctx, force=False, date=None)`rec 含 `(site, kind, date=None)`
- [ ] **Step 8: Commit**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" add inbound_verify/runtime.py
git -C "D:/projects/LogisticsHubIPA/InboundVerify" commit -m "feat(runtime): propagate date through dispatch chain and business-date snapshot" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 7: server.py — TaskRequest.date + 合法性校验 + task_spec
**Files:**
- Modify: `inbound_verify/cli/server.py` — 顶部 import、`TaskRequest``create_task`
**Interfaces:**
- Produces: `POST /tasks` 接受 `{site, kind, force, date}`;合法 date 入队为 `task_spec["date"]`YYYY-MM-DD非法返回 400。
- [ ] **Step 1: 顶部 import 加 timedelta**
`from datetime import datetime``from datetime import datetime, timedelta`
- [ ] **Step 2: TaskRequest 加 date**
```python
class TaskRequest(BaseModel):
site: str
kind: str
force: bool = False
date: Optional[str] = None # YYYY-MM-DD指定则下载该日数据否则走站点 offset
```
- [ ] **Step 3: create_task 加合法性校验 + task_spec 透传 date**
```python
@app.post("/tasks")
def create_task(req: TaskRequest):
"""提交任务 {site, kind, force, date?} → 入队,返回 task_id。"""
if not worker_state["ready"]:
raise HTTPException(
status_code=409, detail="后端尚未就绪,请等待各站点登录完成后再操作"
)
if (req.site, req.kind) not in TASK_HANDLERS:
raise HTTPException(status_code=400, detail=f"无效任务: {req.site}/{req.kind}")
# 指定日期合法性校验(仅在传了 date 时)
if req.date:
try:
target_date = datetime.strptime(req.date, "%Y-%m-%d").date()
except ValueError:
raise HTTPException(
status_code=400, detail=f"date 格式非法,需 YYYY-MM-DD: {req.date}"
)
today = datetime.now().date()
if target_date > today:
raise HTTPException(
status_code=400, detail=f"date 不可为未来日期: {req.date}"
)
if target_date < today - timedelta(days=90):
raise HTTPException(
status_code=400, detail=f"date 超出 90 天回溯上限: {req.date}"
)
if req.site == "百世":
raise HTTPException(
status_code=400, detail="百世固定下载当天,不支持指定日期"
)
task_id = state_store.create_task(req.site, req.kind)
spec = {"site": req.site, "kind": req.kind, "force": req.force}
if req.date:
spec["date"] = req.date
task_queue.put((task_id, spec))
return {"task_id": task_id}
```
- [ ] **Step 4: Black + py_compile**
```bash
PY="D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe"
F="D:/projects/LogisticsHubIPA/InboundVerify/inbound_verify/cli/server.py"
"$PY" -m black "$F" && "$PY" -m py_compile "$F" && echo OK
```
- [ ] **Step 5: Commit**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" add inbound_verify/cli/server.py
git -C "D:/projects/LogisticsHubIPA/InboundVerify" commit -m "feat(server): add date field to POST /tasks with legality validation" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 8: 端到端验证 + 收尾push / 父 bump
**Files:** 无代码改动;验证 + 提交推送。
- [ ] **Step 1: 重启服务加载全部新代码**
停掉旧服务进程重启中通调试模式config.yaml 已是中通):
```bash
PYTHONPATH="D:/projects/LogisticsHubIPA/InboundVerify" "D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe" -m inbound_verify.cli.server
```
(后台运行;等中通登录就绪。)
- [ ] **Step 2: 合法性校验端到端API 400/202**
用 Python urllib避免 curl 中文编码问题)逐一验证,期望:
| 请求 | 期望 |
| --- | --- |
| `{site:"中通", kind:"expected", date:"2026-06-14"}` | 202 + task_id |
| `date:"2099-01-01"`(未来) | 400 "不可为未来日期" |
| `date:"2020-01-01"`(超 90 天) | 400 "超出 90 天回溯上限" |
| `date:"2026/06/14"`(格式错) | 400 "格式非法" |
| `{site:"百世", kind:"undelivered", date:"2026-06-14"}` | 400 "百世…不支持指定日期" |
| `{site:"中通", kind:"expected"}`(不传 date | 202走 offset行为不变 |
```bash
"D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe" - <<'PYEOF'
import json, urllib.request, urllib.error
def post(body):
data = json.dumps(body).encode("utf-8")
req = urllib.request.Request("http://127.0.0.1:8000/tasks", data=data,
headers={"Content-Type":"application/json"}, method="POST")
try:
print(body, "->", urllib.request.urlopen(req, timeout=10).read().decode())
except urllib.error.HTTPError as e:
print(body, "->", e.code, e.read().decode())
post({"site":"中通","kind":"expected","date":"2099-01-01"})
post({"site":"中通","kind":"expected","date":"2020-01-01"})
post({"site":"中通","kind":"expected","date":"2026/06/14"})
post({"site":"百世","kind":"undelivered","date":"2026-06-14"})
PYEOF
```
- [ ] **Step 3: 中通跨月 date 实测(已知 OK**
触发 `{site:"中通", kind:"expected", date:"2026-06-14"}`,观察 worker 日志:出现 `指定日期 2026-06-14折算偏移 …)` + `偏移日期跨月,正在向前翻月导航` + 查询/下载成功。任务 `success`
- [ ] **Step 4: 顺心/韵达/安能 date 实测**
对顺心/韵达/安能各触发一个 expected `date`(取一个近 1 周内的过去日期,确保站点有数据且控件能接受)。观察日志:`指定日期 …` + 正常查询下载。**若某站日期控件不接受字符串而需日历翻月**记录现象按中通同法_zto_flip 模式)追加改动(可能产生新 Task
- [ ] **Step 5: 业务日期快照校验**
下载后 `curl -s http://127.0.0.1:8000/status`,确认对应站 `expected_business_date` == 指定 date。
- [ ] **Step 6: push submodule dev + 父仓库 bump**
```bash
git -C "D:/projects/LogisticsHubIPA/InboundVerify" push origin dev
git -C "D:/projects/LogisticsHubIPA" add InboundVerify
git -C "D:/projects/LogisticsHubIPA" commit -m "chore: bump InboundVerify submodule (date-specific download API)" -m "Co-Authored-By: Claude <noreply@anthropic.com>"
git -C "D:/projects/LogisticsHubIPA" push origin master
```
---
## Self-Reviewplan 作者自检)
1. **Spec 覆盖**接口契约→Task 7透传链路→Task 6中通 effective offset→Task 2韵达/顺心/安能 date→target→Task 3/4/5百世兼容→Task 1业务日期快照→Task 6 Step 5周期调度不动→无需 taskspec 明示合法性校验→Task 7验证→Task 8。✅ 全覆盖。
2. **占位符**:无 TBD/TODO每步含完整代码或精确命令。✅
3. **类型/签名一致**`date=None` 贯穿 server→runtime→sitesactual implyunda/shunxin/anneng原本无 force本计划只加 `date`,与 runtime 调用 `download_func(pg, force=, date=)` / 安能 `lambda(ctx,force,date)` 一致shunxin 双账号 `(pages, foreground, force, date)``_web_handler` 的 list 分支一致。✅
4. **顺序安全**Task 1-5 向后兼容date=None 走 offsetTask 6 启用透传但 server 未传 date仍 NoneTask 7 启用 date。✅

View File

@@ -0,0 +1,332 @@
# 四站点差缺对比逻辑审查报告
> 审查日期2026-07-31
> 审查范围:顺心、中通、韵达、安能 四个站点的应到 vs 实到差缺对比逻辑
> 排除:百世(站点直供未到明细,不参与四站比对)
---
## 一、比对算法总览(四站共用)
`compare.py:process()` 对四个站点执行**完全相同**的算法步骤。站点间的差异仅由 `domain.py:STATIONS` 配置注入——列名映射 + 实到单号解析器。
```
步骤1: 读应到Excel → 按运单号去重keep-first → 构建 {运单号 → (交接单号, 交接件数=n)}
步骤2: 读实到Excel → 站点专用解析器 → 构建 {运单基号 → {已到单号集合}}
步骤3: 逐运单比对
arrived_cnt >= n → 足额到货,跳过
arrived_cnt == 0 → 完全未到
0 < arrived < n → 部分未到
步骤4: 产出未到明细(交接单号 | 运单号 | 总件数 | 已到单号1 | 已到单号2 | ...
```
### 核心口径
| 指标 | 口径 |
|------|------|
| 应到件数 | **交接件数**(非录单件数);按运单号去重 keep-first |
| 实到件数 | 单号去重计数(每扫描一件=一个单号) |
| 未到件数 | max(0, 应到件数 实到件数) |
| 未到率 | 未到件数 ÷ 应到件数 |
### 未到明细输出约定
- 仅列出**短少运单**(实到 < 应到)
- 列出该运单**实际已到的单号**已到单号1, 已到单号2, ...
- **不编造缺件子单号**——实到扫描顺序号乱序,无法反推缺了哪个顺序号
### 统计指标
| 指标 | 含义 |
|------|------|
| 运单数 | 应到运单去重数 |
| 应到件 | Σ 交接件数 |
| 已到件 | Σ 实到单号去重数 |
| 未到件 | max(0, 应到件 已到件) |
| 涉及运单 | 存在短少的运单数 |
| 完全未到 | 整单零到货运单数 |
| 部分未到 | 部分缺件运单数 |
---
## 二、四站配置对照
`domain.py:STATIONS` — 所有差异集中于此配置表,比对核心代码不感知站点差异。
| 维度 | 中通 | 顺心 | 韵达 | 安能 |
|------|------|------|------|------|
| 应到文件 | `中通-应到货物数据.xlsx` | `顺心-应到货物数据.xlsx` | `韵达-应到货物数据.xlsx` | `安能-应到货物数据.xlsx` |
| 实到文件 | `中通-实到货物数据.xlsx` | `顺心-实到货物数据.xlsx` | `韵达-实到货物数据.xlsx` | `安能-实到货物数据.xlsx` |
| 应到-运单号列 | `运单号` | `运单号` | `运单号` | `运单号` |
| 应到-件数列 | `交接件数` | `交接件数` | `交接件数` | `交接件数` |
| 应到-交接单号列 | `交接单号` | `交接单号` | `交接单号` | `交接单号` |
| 实到-基号列 | —(从复合串推导) | `运单号` | **`主单号`** | **`所属单号`** |
| 实到-单号列 | `运单号`(复合串) | `子单号` | `子单号` | `扫描单号` |
| 解析器 | `arrived_pieces_zhongtong` | `arrived_pieces_by_cols` | `arrived_pieces_by_cols` | `arrived_pieces_by_cols` |
---
## 三、逐站点详细分析
### 3.1 中通ZTO
#### 业务逻辑
实到货物数据中的「运单号」为复合串,由三部分构成:
```
┌──────────┬────────────┬──────────┐
│ 运单号 │ 录单件数 │ 顺序号 │
│ (12位) │ (4位) │ (4位) │
└──────────┴────────────┴──────────┘
总长 20 位
示例: 330953527953 0001 0001
├─ 运单号 ─┤├录单┤├顺序┤
```
- **运单号(12位)**: 与应到货物数据中的运单号对齐
- **录单件数(4位)**: 该运单在系统中的录单总件数0占位
- **顺序号(4位)**: 0占位`0001`, `0002`, `0003`, `0004`
对比逻辑:
1. 从应到数据取运单号 + 交接件数(**非录单件数**
2. 从实到数据取复合串掐尾8位得运单基号完整串为子运单号
3. 按运单基号分组,子运单号去重得实到件数
4. 实到件数 < 交接件数 → 差缺
> **重要**: 录单件数仅作参考。举例:某运单录单件数=4、交接件数=2实到最多出现2条数据。如果只出现了1条我们只知道差缺了但**无法判断具体差缺了哪一件**(顺序号乱序)。
#### 代码实现
`domain.py:17-26` — 实到解析器:
```python
def arrived_pieces_zhongtong(df):
res = defaultdict(set)
for v in df["运单号"]:
v = str(v).strip()
if len(v) > 8 and v[-4:].isdigit():
res[v[:-8]].add(v) # 基号=前12位, 已到单号=完整20位复合串
return res
```
`domain.py:48-56` — 站点配置:
```python
{
"name": "中通",
"exp_qty": "交接件数", # 应到件数口径:交接件数(非录单件数)
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_zhongtong,
"columns": ["交接单号", "运单号", "总件数"],
}
```
#### 对齐情况:✅ 对齐
代码实现与业务逻辑一致。`v[:-8]` 掐尾8位得12位运单基号保留完整复合串作为已到单号——不解析、不推断录单件数和顺序号的具体含义。
---
### 3.2 安能Anneng
#### 业务逻辑
与中通相同的差缺对比逻辑。
安能实到数据同样为复合串,结构:`运单号(12位) + 录单件数(4位) + 顺序号(4位)`20位
与中通的关键区别:安能实到表有**独立的「所属单号」列**干净运单基号无需像中通那样从复合串掐尾8位推导基号。
#### 代码实现
`domain.py:76-86`
```python
{
"name": "安能",
"arrived_pieces": arrived_pieces_by_cols("所属单号", "扫描单号"),
...
}
```
安能使用 `arrived_pieces_by_cols` 而非 `arrived_pieces_zhongtong`——直接从「所属单号」列读基号、从「扫描单号」列读完整单号,效果等价。
| 差异点 | 中通 | 安能 |
|--------|------|------|
| 实到基号来源 | 从复合串解析(`v[:-8]` | 直接读「所属单号」列 |
| 实到单号来源 | 复合串本身(「运单号」列) | 「扫描单号」列 |
| 解析器 | `arrived_pieces_zhongtong` | `arrived_pieces_by_cols` |
| 最终产出 | `{基号 → {完整单号集合}}` | 相同 |
#### 数据库验证
```
piece_no=61003282264500140014 → waybill_no=610032822645 (12位), total=0014, seq=0014
```
#### 对齐情况:✅ 对齐
---
### 3.3 顺心Shunxin
#### 业务逻辑
顺心站点需区分两类运单:
**A. 非SF开头运单占 97%:**
实到「子单号」结构为两部分:
```
┌──────────┬──────────┐
│ 运单号 │ 顺序号 │
│ (不定长) │ (3位) │
└──────────┴──────────┘
示例: S71623721115 001
├─ 运单号 ──┤├顺序┤
注意:顺心子单号无录单件数部分(仅两部分)
```
对比时从实到取「子单号」列,按「运单号」分组,子单号去重得实到件数。
**B. SF开头运单占 3%:**
SF订单的「子单号」为**随机号码**(非由运单号衍生),不能用于差缺推导。
对比逻辑:
1. 在实到数据中按「运单号」字段查找,统计出现次数
2. 出现次数 < 交接件数 → 差缺
3. 将找到的子单号(虽随机但可以列出来)填入「已到单号」列
SF订单的差缺判定**只基于交接件数与实到运单号出现次数的比较**,不依赖子单号的结构解析。
#### 代码实现
`domain.py:57-65`
```python
{
"name": "顺心",
"arrived_pieces": arrived_pieces_by_cols("运单号", "子单号"),
}
```
**SF 与非 SF 没有任何区分处理。** 所有运单走同一条路径。
#### 数据库验证
**非SF正常:**
```
子单号=S71623721115001 → 运单号=S71623721115 + 后缀=001 ✅
子单号=S71934073996002 → 运单号=S71934073996 + 后缀=002 ✅
```
**SF异常:**
```
运单号=SF1225002296515 的两条实到记录:
子单号=SF2025318183224 (随机SF号码)
子单号=SF1225002296515 (与运单号相同)
```
数据中有 10 个SF运单存在多条实到记录。
#### 对齐情况:⚠️ 部分对齐SF特殊逻辑缺失
| 检查项 | 代码现状 | 业务要求 |
|--------|----------|----------|
| 非SF处理 | ✅ `arrived_pieces_by_cols("运单号", "子单号")` | 一致 |
| 非SF子单号结构 | ✅ 运单号 + 顺序号(两部分) | 一致 |
| SF处理 | ❌ 与非SF完全一致使用子单号去重 | **不能**使用子单号,只按运单号行数计数 |
| 功能影响 | 子单号虽随机但值唯一,按目前逻辑也能正确去重计数 | 但语义不正确——SF子单号不由运单号衍生 |
---
### 3.4 韵达Yunda
#### 业务逻辑
**去重规则:** 韵达实到数据存在重复行(同一子单号出现两次)。去重依据为「交接单号」字段:
- **保留**交接单号为**空**的行
- **丢弃**交接单号**非空**的行
**子单号结构:** 两部分——单号 + 顺序号(无录单件数部分)。
```
┌──────────┬──────────┐
│ 主单号 │ 顺序号 │
│ (不定长) │ (4位) │
└──────────┴──────────┘
示例: 713326603 0003
├─主单号─┤├顺序┤
```
**对比方式:** 与中通/安能同——按「主单号」分组,「子单号」去重得实到件数,与交接件数比对。
#### 代码实现
`store.py:316-320`(入库过滤):
```python
if site == "韵达":
# 韵达业务清洗:抛弃「交接单号」为空的行(派件/签收等其他扫描无交接单号),
# 再按子单号去重一件多扫只留一条清洗后子单号已天然唯一drop 为保险)。
df = df[df["交接单号"].astype(str).str.strip() != ""] # ← 保留非空
df = df.drop_duplicates(subset=[cm["piece"]], keep="last")
```
`domain.py:67-75`(比对配置):
```python
{
"name": "韵达",
"exp_wb": "运单号",
"arrived_pieces": arrived_pieces_by_cols("主单号", "子单号"),
}
```
#### 对齐情况:❌ 交接单号过滤逻辑完全相反
| 检查项 | 代码现状 | 业务要求 |
|--------|----------|----------|
| 交接单号过滤 | 保留 `!= ""`**非空** | 保留 `== ""`**空** |
| 子单号结构 | ✅ `7133266030003` = wb`713326603` + seq`0003` | 一致 |
| 实到解析 | ✅ `arrived_pieces_by_cols("主单号", "子单号")` | 一致 |
| compare.py 过滤 | ❌ **无过滤**,所有行参与比对 | 需要过滤 |
**影响分析:**
1. `store.py` 过滤反了——入库时留下了错误的数据集
2. `compare.py` 完全没有交接单号过滤——如果原始 Excel 中同时存在空和非空行,比对阶段会全部读入导致重复计数
3. 当前数据库中韵达 3483 条记录全部为非空交接单号——说明当前 Excel 数据中空交接单号行偏少或不存在,但这不改变逻辑错误
---
## 四、差异汇总
| # | 站点 | 问题 | 严重程度 | 影响范围 |
|---|------|------|----------|----------|
| 1 | **韵达** | 交接单号过滤反了:`!= ""` 应改为 `== ""` | ❌ 严重 | `store.py:319` + `compare.py` 需新增过滤 |
| 2 | **顺心** | SF运单无特殊处理与非SF混用子单号 | ⚠️ 中等 | `domain.py` 需新增SF判断分支 |
| 3 | **中通** | 录单件数0占位描述与实际数据完全一致 | ✅ 无影响 | 代码不依赖此区分 |
---
## 五、代码位置索引
| 逻辑 | 文件 | 行号 |
|------|------|------|
| 单站比对 `process()` | `compare.py` | 61-143 |
| 站点配置 `STATIONS` | `domain.py` | 46-87 |
| 中通实到解析器 | `domain.py` | 17-26 |
| 通用实到解析器 | `domain.py` | 29-42 |
| 单站未到文件写入 | `compare.py` | 258-273 |
| 全量汇总报告 | `compare.py` | 293-324 |
| 未到触发编排 | `runtime.py` | 478-497 |
| 韵达入库过滤(需修) | `store.py` | 316-320 |
| 顺心实到配置(需修) | `domain.py` | 57-65 |

View File

@@ -0,0 +1,264 @@
# 顺心 DB 差缺对比 — 实施计划
> 日期2026-07-31
> 目标:将顺心站点差缺对比从 Excel 读取改为 PostgreSQL 查询,并修正 SF 运单特殊处理逻辑
---
## 一、背景
### 当前状态Excel 方式)
```
compare.py:process("顺心")
├── 读 downloads/顺心-应到货物数据.xlsx
├── 读 downloads/顺心-实到货物数据.xlsx
├── arrived_pieces_by_cols("运单号", "子单号") ← SF/non-SF 无区分
└── 产出 {站}-未到数据.xlsx + 统计 dict
```
### 需要解决的两个问题
1. **从 Excel 切换到 DB**:数据已持久化到 PostgreSQL比对应直接从 DB 查询
2. **顺心 SF 运单特殊处理**SF 运单的子单号为随机号码,不能用于去重计数,应使用行计数
---
## 二、数据结构
### PostgreSQL 表
**expected_record**(关键列):
| 列 | 类型 | 说明 |
|----|------|------|
| site | TEXT | 站点 |
| waybill_no | TEXT | 运单号唯一键之一SF 以 "SF" 开头) |
| handover_no | TEXT | 交接单号(批次标识) |
| handover_pieces | INTEGER | 交接件数(应到口径) |
| order_pieces | INTEGER | 录单件数(参考) |
| business_date | DATE | 下载目标日期 |
**actual_record**(关键列):
| 列 | 类型 | 说明 |
|----|------|------|
| site | TEXT | 站点 |
| waybill_no | TEXT | 运单基号(关联 expected_record |
| piece_no | TEXT | 扫描单号non-SF运单号+顺序号SF随机号码 |
| scan_time | TIMESTAMPTZ | 扫描时间(可靠,当天数据=当天扫描) |
### SF 数据特征(已验证)
- 顺心 actual_record 中 SF 运单148 条
- `piece_no == waybill_no`86 条58%
- `piece_no != waybill_no`62 条42%)← 随机 SF 号码
- SF 运单 expected99 条,分布在 31 个交接批次中
---
## 三、算法设计
### 核心思路:以实到为锚,通过交接单号反推批次
```
输入: site="顺心", date="2026-07-25"
Step 1 — 取实到锚点
SELECT DISTINCT waybill_no FROM actual_record
WHERE site='顺心' AND scan_time::date = '2026-07-25'
Step 2 — 反推交接批次
SELECT DISTINCT handover_no FROM expected_record
WHERE site='顺心'
AND waybill_no IN (Step 1 的运单集合)
Step 3 — 展开批次全量应到
SELECT waybill_no, handover_no, handover_pieces
FROM expected_record
WHERE site='顺心'
AND handover_no IN (Step 2 的交接单号集合)
Step 4 — 取批次全量实到
SELECT waybill_no, piece_no FROM actual_record
WHERE site='顺心'
AND waybill_no IN (Step 3 的运单集合)
Step 5 — 逐运单比对
for each waybill in Step 3:
if waybill_no LIKE 'SF%':
arrived_cnt = COUNT(*) ← 行计数,不去重
else:
arrived_cnt = COUNT(DISTINCT piece_no) ← 子单号去重
if arrived_cnt < handover_pieces → 差缺
```
### SF vs non-SF 处理差异
| | non-SF | SF |
|------|--------|-----|
| piece_no 含义 | 运单号 + 顺序号(可推导) | 随机 SF 号码(无推导意义) |
| 实到计数方式 | `COUNT(DISTINCT piece_no)` | `COUNT(*)`(行计数) |
| 已到单号列表 | 列出去重后的子单号 | 列出所有 piece_no含重复 |
### 统计指标
| 指标 | 公式 |
|------|------|
| 运单数 | Step 3 去重运单数 |
| 应到件 | Σ handover_pieces |
| 已到件 | Σ arrived_cnt |
| 未到件 | max(0, 应到件 已到件) |
| 涉及运单 | arrived_cnt < handover_pieces 的运单数 |
| 完全未到 | arrived_cnt = 0 的运单数 |
| 部分未到 | 0 < arrived_cnt < handover_pieces 的运单数 |
| 未到率 | 未到件 ÷ 应到件 |
### 边界情况覆盖
| 情况 | 覆盖方式 |
|------|----------|
| 同日多批次 | Step 2 查出全部涉及的 handover_no |
| 跨天到达(延迟) | Step 4 不限 scan_time历史扫描全计入 |
| 溢到(实到 > 应到) | arrived_cnt >= n 跳过,不进差缺表 |
| 完全沉默批次 | 一件未扫 = 实到无锚点,该批次不会被触发——在首次有扫描那天被纳入 |
| SF 子单号重复 | 用 COUNT(*) 而非 COUNT(DISTINCT),不会漏计 |
---
## 四、模块设计
### 新增文件
**`inbound_verify/db_compare.py`** — DB 比对引擎(纯 PostgreSQL + Python
```python
# 核心函数签名
def compare_site_date(site: str, date: str) -> CompareResult | None:
"""对指定站点和日期执行 DB 差缺比对。
返回 CompareResultstats + undelivered_rows
当天无实到数据时返回 None。
"""
def compare_site_batch(site: str, handover_no: str) -> CompareResult | None:
"""按指定交接单号执行全批次比对(不依赖实到锚点)。"""
```
**数据类型**
```python
@dataclass
class CompareResult:
stats: dict # 统计指标
rows: list[dict] # 差缺明细行
batches: list[str] # 涉及的交接批次
@dataclass
class UndeliveredRow:
handover_no: str # 交接单号
waybill_no: str # 运单号
total_pieces: int # 总件数(=交接件数)
arrived_pieces: int # 已到件数
arrived_list: list[str] # 已到单号列表
is_sf: bool # 是否 SF 运单
```
### 修改文件
**`inbound_verify/cli/server.py`** — 新增 API 端点
```python
@app.post("/compare")
def run_compare(req: CompareRequest):
"""DB 比对:{site, date} → 返回差缺结果"""
@app.get("/compare/{site}/{date}")
def get_compare(site: str, date: str):
"""查询某站点某日的差缺结果(缓存)"""
```
### 现有文件保持不动
- `compare.py` — 保留不动Excel 比对继续可用
- `domain.py` — 可能需要新增 DB 版站点配置(或复用现有)
- `runtime.py` — 暂不改动,`_site_undelivered_handler` 仍走 Excel 路径
---
## 五、实施步骤
### Phase 1 — `db_compare.py` 核心引擎
- [ ] 新建 `inbound_verify/db_compare.py`
- [ ] 实现 `compare_site_date("顺心", date)`
- [ ] SF/non-SF 分支处理
- [ ] 返回 `CompareResult`
- [ ] 终端手动验证(直接调函数,打印结果)
### Phase 2 — API 端点
- [ ]`server.py` 新增 `POST /compare`
- [ ] `CompareRequest { site, date }`
- [ ] 调用 `db_compare.compare_site_date()`
- [ ] 返回 JSONstats + undelivered rows
- [ ] HTTP 验证curl 调 `/compare` 对比不同日期结果
### Phase 3 — Excel 输出(可选)
- [ ] `db_compare` 生成 Excel 报告(复用现有 `compare.py` 的 openpyxl 样式)
- [ ] 输出到 `output/顺心-{date}-未到数据.xlsx`
- [ ] 或者只输出 JSON前端自行渲染
### Phase 4 — 替换 undelivered 任务流
- [ ] `runtime.py` 新增 `_db_undelivered_handler`
- [ ] 下载完成后不再调 Excel 比对,改调 DB 比对
- [ ] 逐步替换 `TASK_HANDLERS` 中的顺心 undelivered handler
### Phase 5 — 扩展到中通/韵达/安能
- [ ] 各站适配(主要是 piece_no 去重方式差异)
- [ ] 中通:`COUNT(DISTINCT piece_no)`,无 SF 问题
- [ ] 韵达:同上
- [ ] 安能:同上
---
## 六、测试策略
### 手工验证Phase 1
```python
# 终端直接调
from inbound_verify.db_compare import compare_site_date
result = compare_site_date("顺心", "2026-07-25")
print(result.stats)
# 对比基于 Excel 版的 compare.process("顺心") 结果
```
### API 验证Phase 2
```bash
curl -X POST http://127.0.0.1:8000/compare \
-H "Content-Type: application/json" \
-d '{"site":"顺心","date":"2026-07-25"}'
```
### 回归验证
- 新 DB 比对结果 vs 旧 Excel 比对结果(同一份数据)
- SF 运单的 arrived_cnt 对比DB 版COUNT(*)vs Excel 版COUNT DISTINCT piece_no
- 确认 SF 运单不再被漏计
---
## 七、风险与注意事项
| 风险 | 缓解 |
|------|------|
| DB 连接超时cpolar 隧道) | 加 connect_timeout + try/except 降级 |
| 全表扫描性能 | 依赖 (site, waybill_no) 和 (site, scan_time) 索引 |
| SF 运单数据量小(~1% | 测试覆盖可能不足——需找有 SF 差缺的日期验证 |
| `scan_time` 时区 | 统一用 `::date` cast确认与服务器时区一致 |

View File

@@ -0,0 +1,579 @@
# Tier 1: Package Move Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Move all 12 flat root-level Python modules into an installable `inbound_verify/` package (sites/, cli/ subpackages), rewrite every internal import to package-qualified form, and expose three console_script entry points — with **zero behavior change**.
**Architecture:** Pure mechanical relocation (git mv preserves history) + import rewrite + paths.py anchor fix + entry-point `main()` wrappers + `pyproject.toml` packaging. No logic changes. Module names kept for heavily-referenced modules (`state_store`, `expected_undelivered`, `runtime`, `paths`) to avoid call-site churn; only leaf entries (`db_store→store`, `main_router→cli/router`, `server→cli/server`) and site files (drop `site_` prefix) are renamed.
**Tech Stack:** Python ≥3.10, setuptools (PEP 517/621), pip editable install, psycopg3, FastAPI/uvicorn, Playwright.
## Global Constraints
- **Python ≥ 3.10** (`requires-python = ">=3.10"` in pyproject).
- **Package** import name `inbound_verify`; **distribution** name `inbound-verify`.
- **Dependency floors** (verbatim from spec): `pandas>=2.0.0`, `playwright>=1.40.0`, `openpyxl>=3.1.0`, `PyYAML>=6.0`, `websocket-client>=1.0.0`, `fastapi>=0.110.0`, `uvicorn>=0.27.0`, `apscheduler>=3.10.0`, `psycopg[binary]>=3.1`.
- **No test suite** (user decision). Verification = `compileall` + import smoke + grep-for-stale-refs + DB connectivity. No pytest.
- **Behavior must not change** in Tier 1 — pure move.
- **No auto-commit/push.** Every commit step below runs ONLY after the user explicitly says "提交/commit". Commit messages in English.
- **Black-format** every changed `.py` (global rule).
- All changes are **inside the `InboundVerify` git submodule**; the parent repo pointer bump is a separate parent-repo step, out of scope.
- All commands run from the `InboundVerify/` directory using the venv interpreter `.venv/Scripts/python.exe` (Windows; no activation needed).
**Reference spec:** `docs/superpowers/specs/2026-07-23-package-restructure-design.md` (§3 mapping table, §4 paths anchor, §5 import rules, §6 entry/packaging, §7 verification gate).
---
## File Structure (what each file becomes responsible for)
```
inbound_verify/
├── __init__.py # empty (package marker)
├── paths.py # path anchors → PROJECT ROOT (one dir above package)
├── runtime.py # orchestration core (unchanged logic)
├── state_store.py # SQLite state (unchanged; name kept)
├── expected_undelivered.py# offline compare (unchanged; name kept — Tier 2 renames to compare)
├── store.py # PostgreSQL persist + main() (was db_store.py)
├── sites/
│ ├── __init__.py # empty
│ ├── shunxin.py # (was site_shunxin.py)
│ ├── baishi.py # (was site_baishi.py)
│ ├── zto.py # (was site_zto.py)
│ ├── yunda.py # (was site_yunda.py)
│ └── anneng.py # (was site_anneng.py)
└── cli/
├── __init__.py # empty
├── router.py # interactive menu + main() (was main_router.py)
└── server.py # FastAPI service + main() (was server.py)
```
Root keeps: `pyproject.toml` (new), `config.yaml`, `config.example.yaml`, `schema.sql`, `requirements.txt`, `README.md`, `CLAUDE.md`, `docs/`, `downloads/`, `output/`, `state/`.
---
## Task 1: Package scaffold + pyproject + editable install
**Files:**
- Create: `inbound_verify/__init__.py`, `inbound_verify/sites/__init__.py`, `inbound_verify/cli/__init__.py`
- Create: `pyproject.toml`
**Interfaces:**
- Produces: an importable (near-empty) `inbound_verify` package + console_script registration. The flat root scripts remain 100% functional after this task (untouched).
- [ ] **Step 1: Create package marker files**
Create three empty files:
- `inbound_verify/__init__.py`
- `inbound_verify/sites/__init__.py`
- `inbound_verify/cli/__init__.py`
Each is a single comment line:
```python
# inbound_verify package
```
- [ ] **Step 2: Write pyproject.toml**
Create `pyproject.toml`:
```toml
[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"
[project]
name = "inbound-verify"
version = "0.1.0"
description = "物流到货数据自动下载与应到未到核对工具"
requires-python = ">=3.10"
dependencies = [
"pandas>=2.0.0",
"playwright>=1.40.0",
"openpyxl>=3.1.0",
"PyYAML>=6.0",
"websocket-client>=1.0.0",
"fastapi>=0.110.0",
"uvicorn>=0.27.0",
"apscheduler>=3.10.0",
"psycopg[binary]>=3.1",
]
[project.scripts]
inbound-verify = "inbound_verify.cli.router:main"
inbound-verify-server = "inbound_verify.cli.server:main"
inbound-verify-db = "inbound_verify.store:main"
[tool.setuptools.packages.find]
include = ["inbound_verify*"]
```
- [ ] **Step 3: Editable-install into the existing venv**
Run:
```bash
.venv/Scripts/python.exe -m pip install -e .
```
Expected: `Successfully installed inbound-verify-0.1.0` (deps already satisfied from earlier install — no network needed).
- [ ] **Step 4: Verify package imports**
Run:
```bash
.venv/Scripts/python.exe -c "import inbound_verify, inbound_verify.sites, inbound_verify.cli; print('package OK')"
```
Expected output: `package OK`
- [ ] **Step 5: Commit (only after user confirms)**
```bash
git add inbound_verify/__init__.py inbound_verify/sites/__init__.py inbound_verify/cli/__init__.py pyproject.toml
git commit -m "chore: scaffold inbound_verify package and pyproject
Co-Authored-By: Claude <noreply@anthropic.com>"
```
> Do NOT commit until the user says to.
---
## Task 2: Atomic move — relocate, rewrite imports, fix anchor, wire mains, verify
**Files:**
- Move (git mv): all 12 root `.py` modules → package locations (see Step 1)
- Modify: `inbound_verify/paths.py` (anchor), and import lines + call sites in every moved module
- Modify: entry `main()` in `cli/router.py`, `cli/server.py`, `store.py`
**Interfaces:**
- Consumes: the scaffold from Task 1 (package importable + editable install active).
- Produces: a fully functional `inbound_verify` package invokable via `python -m inbound_verify.cli.router`, `python -m inbound_verify.cli.server`, `python -m inbound_verify.store`, or the three console_scripts. Old root `.py` files are gone. The flat root scripts no longer exist — invocation switches to package form.
> **Why this is one task:** in a flat-import codebase, moving `paths.py` (imported by everyone) immediately breaks every importer until ALL moves + rewrites are complete. There is no intermediate state that imports cleanly, so the whole move is one atomic unit verified by the gate at the end. Each file-edit step below is followed by `py_compile` of that file to catch syntax errors as we go.
- [ ] **Step 1: Relocate all 12 modules with git mv (history preserved)**
From the `InboundVerify/` directory:
```bash
git mv paths.py inbound_verify/paths.py
git mv runtime.py inbound_verify/runtime.py
git mv state_store.py inbound_verify/state_store.py
git mv expected_undelivered.py inbound_verify/expected_undelivered.py
git mv db_store.py inbound_verify/store.py
git mv main_router.py inbound_verify/cli/router.py
git mv server.py inbound_verify/cli/server.py
git mv site_shunxin.py inbound_verify/sites/shunxin.py
git mv site_baishi.py inbound_verify/sites/baishi.py
git mv site_zto.py inbound_verify/sites/zto.py
git mv site_yunda.py inbound_verify/sites/yunda.py
git mv site_anneng.py inbound_verify/sites/anneng.py
```
After this the tree is temporarily broken (imports unresolved) — expected. Continue.
- [ ] **Step 2: Fix paths.py anchor to point at project root**
In `inbound_verify/paths.py`, replace:
```python
BASE_DIR = os.path.dirname(os.path.abspath(__file__))
```
with:
```python
# __file__ = <root>/inbound_verify/paths.py → 上两级 = 项目根
BASE_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
```
(`DOWNLOAD_DIR`/`OUTPUT_DIR`/`CONFIG_PATH`/`STATE_DB_PATH` lines stay unchanged — they derive from BASE_DIR.)
Verify syntax:
```bash
.venv/Scripts/python.exe -m py_compile inbound_verify/paths.py
```
Expected: no output (success).
- [ ] **Step 3: Rewrite imports in state_store.py**
In `inbound_verify/state_store.py`, replace:
```python
from paths import STATE_DB_PATH
```
with:
```python
from inbound_verify.paths import STATE_DB_PATH
```
Verify: `.venv/Scripts/python.exe -m py_compile inbound_verify/state_store.py` → no output.
- [ ] **Step 4: Rewrite imports + call sites in runtime.py**
In `inbound_verify/runtime.py`, replace the import block:
```python
from paths import DOWNLOAD_DIR, CONFIG_PATH
import state_store
import site_shunxin
import site_baishi
import site_zto
import site_yunda
import site_anneng
import expected_undelivered # dispatch 的 compare 任务用
```
with:
```python
from inbound_verify.paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify import state_store
from inbound_verify.sites import shunxin, baishi, zto, yunda, anneng
from inbound_verify import expected_undelivered # dispatch 的 compare 任务用
```
Then **drop the `site_` prefix at every call site** (28 references — 18 in this file). Apply these 5 replacements (all occurrences):
- `site_shunxin.``shunxin.`
- `site_baishi.``baishi.`
- `site_zto.``zto.`
- `site_yunda.``yunda.`
- `site_anneng.``anneng.`
Affected lines in runtime.py (for reference, all must change): 36, 37, 38, 39, 118, 139, 336, 408, 417, 463, 464, 467, 469, 470, 472, 473, 475, 476.
`state_store.` and `expected_undelivered.` call sites stay UNCHANGED (names kept). Confirm the key handler block now reads:
```python
TASK_HANDLERS = {
("顺心", "expected"): _web_handler("顺心", shunxin.shunxin_expected_download),
("顺心", "actual"): _web_handler("顺心", shunxin.shunxin_actual_download),
("顺心", "undelivered"): _site_undelivered_handler("顺心"),
("百世", "undelivered"): _web_handler(
"百世", baishi.baishi_download_undelivered_data
),
("中通", "expected"): _web_handler("中通", zto.zto_expected_download),
("中通", "actual"): _web_handler("中通", zto.zto_actual_download),
("中通", "undelivered"): _site_undelivered_handler("中通"),
("韵达", "expected"): _web_handler("韵达", yunda.yunda_expected_download),
("韵达", "actual"): _web_handler("韵达", yunda.yunda_actual_download),
("韵达", "undelivered"): _site_undelivered_handler("韵达"),
("安能", "expected"): lambda ctx: anneng.anneng_expected_download(),
("安能", "actual"): lambda ctx: anneng.anneng_actual_download(),
("安能", "undelivered"): _site_undelivered_handler("安能"),
("__compare__", "compare"): lambda ctx: (expected_undelivered.main() or True),
}
```
Verify: `.venv/Scripts/python.exe -m py_compile inbound_verify/runtime.py` → no output.
- [ ] **Step 5: Rewrite imports in expected_undelivered.py**
This module has TWO lazy `import state_store` statements (inside functions). Replace each occurrence of:
```python
import state_store
```
with:
```python
from inbound_verify import state_store
```
(There is no `from paths import` here — the module has its own `BASE/DOWNLOADS/OUTPUT` constants; that dedup is Tier 2, not now.)
Verify: `.venv/Scripts/python.exe -m py_compile inbound_verify/expected_undelivered.py` → no output.
- [ ] **Step 6: Rewrite imports + rename _cli→main in store.py**
In `inbound_verify/store.py`, replace:
```python
from paths import BASE_DIR, CONFIG_PATH, DOWNLOAD_DIR
```
with:
```python
from inbound_verify.paths import BASE_DIR, CONFIG_PATH, DOWNLOAD_DIR
```
Replace:
```python
import expected_undelivered as eu # 复用站点 / 文件名 / 列映射 / 基号口径(单一来源)
```
with:
```python
from inbound_verify import expected_undelivered as eu # 复用站点 / 文件名 / 列映射 / 基号口径(单一来源)
```
(`eu.` call sites stay unchanged.)
Rename the CLI entry: replace the function definition:
```python
def _cli():
```
with:
```python
def main():
```
And at the bottom replace:
```python
if __name__ == "__main__":
_cli()
```
with:
```python
if __name__ == "__main__":
main()
```
Verify: `.venv/Scripts/python.exe -m py_compile inbound_verify/store.py` → no output.
- [ ] **Step 7: Rewrite imports in all 5 site modules**
In each of `inbound_verify/sites/{shunxin,baishi,zto,yunda,anneng}.py`, replace:
```python
from paths import DOWNLOAD_DIR, CONFIG_PATH
import state_store
```
with:
```python
from inbound_verify.paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify import state_store
```
(Each site module has exactly these two internal imports; `state_store.` call sites unchanged.)
Verify all five:
```bash
.venv/Scripts/python.exe -m py_compile inbound_verify/sites/shunxin.py inbound_verify/sites/baishi.py inbound_verify/sites/zto.py inbound_verify/sites/yunda.py inbound_verify/sites/anneng.py
```
Expected: no output.
- [ ] **Step 8: Rewrite imports + call sites + add main() in cli/router.py**
In `inbound_verify/cli/router.py`, replace the import block (lines ~1431):
```python
from paths import CONFIG_PATH
from runtime import (
APP_SITES,
HEARTBEAT_INTERVAL,
dispatch_task,
launch_and_prepare,
run_heartbeat,
)
import state_store
# 各站点模块(自动化测试 + 比对用;任务派发在 runtime
import site_shunxin
import site_baishi
import site_zto
import site_yunda
import site_anneng
import expected_undelivered
```
with:
```python
from inbound_verify.paths import CONFIG_PATH
from inbound_verify.runtime import (
APP_SITES,
HEARTBEAT_INTERVAL,
dispatch_task,
launch_and_prepare,
run_heartbeat,
)
from inbound_verify import state_store
# 各站点模块(自动化测试 + 比对用;任务派发在 runtime
from inbound_verify.sites import shunxin, baishi, zto, yunda, anneng
from inbound_verify import expected_undelivered
```
Drop the `site_` prefix at the 10 call sites (apply the same 5 replacements as Step 4). Affected lines: 60, 61, 64, 65, 68, 69, 72, 73, 86, 105. (`expected_undelivered.main()` at line 38 stays unchanged.)
Add an entry function and update the `__main__` guard. Replace:
```python
if __name__ == "__main__":
run_multi_site_daemon()
```
with:
```python
def main():
"""交互菜单模式入口。"""
run_multi_site_daemon()
if __name__ == "__main__":
main()
```
Verify: `.venv/Scripts/python.exe -m py_compile inbound_verify/cli/router.py` → no output.
- [ ] **Step 9: Rewrite imports + add main() in cli/server.py**
In `inbound_verify/cli/server.py`, replace:
```python
from paths import DOWNLOAD_DIR, OUTPUT_DIR
import state_store
from runtime import (
HEARTBEAT_INTERVAL,
TASK_HANDLERS,
dispatch_task,
launch_and_prepare,
run_heartbeat,
)
```
with:
```python
from inbound_verify.paths import DOWNLOAD_DIR, OUTPUT_DIR
from inbound_verify import state_store
from inbound_verify.runtime import (
HEARTBEAT_INTERVAL,
TASK_HANDLERS,
dispatch_task,
launch_and_prepare,
run_heartbeat,
)
```
Replace the bottom entry block:
```python
if __name__ == "__main__":
uvicorn.run(app, host="0.0.0.0", port=8000)
```
with:
```python
def main():
"""服务模式入口。传字符串导入路径(规范写法;不开 reload/workers 时进程内 import行为等价"""
uvicorn.run("inbound_verify.cli.server:app", host="0.0.0.0", port=8000)
if __name__ == "__main__":
main()
```
Verify: `.venv/Scripts/python.exe -m py_compile inbound_verify/cli/server.py` → no output.
- [ ] **Step 10: Clean up stale filename comments**
Cosmetic but keeps grep clean (Step 12 depends on it). In each `inbound_verify/sites/*.py`, update the line-1 header `# site_xxx.py``# sites/xxx.py`. In `inbound_verify/sites/anneng.py`, update the standalone-run comment near the top:
```python
# .venv/Scripts/python.exe site_anneng.py
```
```python
# python -m inbound_verify.sites.anneng expected # 或 actual
```
And the comment at the `CDP_PORT` line referencing "独立运行 site_anneng.py" — update to "独立运行python -m inbound_verify.sites.anneng".
- [ ] **Step 11: Black-format all changed files**
```bash
.venv/Scripts/python.exe -m black inbound_verify
```
Expected: `reformatted ...` / `left unchanged` lines, exit 0.
- [ ] **Step 12: VERIFICATION GATE — run all five checks**
**12a. compileall (syntax across whole package):**
```bash
.venv/Scripts/python.exe -m compileall inbound_verify
```
Expected: no errors.
**12b. Import smoke (catches every wrong import path / missed rewrite):**
```bash
.venv/Scripts/python.exe -c "import inbound_verify.cli.router, inbound_verify.cli.server, inbound_verify.store, inbound_verify.runtime, inbound_verify.state_store; print('import smoke OK')"
```
Expected: `import smoke OK`. (The three entries transitively import sites + expected_undelivered.)
**12c. No stale `site_` references:**
```bash
grep -rn "site_shunxin\|site_baishi\|site_zto\|site_yunda\|site_anneng" inbound_verify || echo "no stale site_ refs OK"
```
Expected: `no stale site_ refs OK`.
**12d. No stale bare flat imports:**
```bash
grep -rnE "^from paths import|^from runtime import|^import site_|^import state_store$|^import expected_undelivered$" inbound_verify || echo "no stale flat imports OK"
```
Expected: `no stale flat imports OK`.
**12e. paths anchor points at project root + DB still connects:**
```bash
.venv/Scripts/python.exe -c "from inbound_verify.paths import BASE_DIR; print('BASE_DIR', BASE_DIR)"
.venv/Scripts/python.exe -c "from inbound_verify.store import _connect, _load_pg_config; c=_load_pg_config(); conn=_connect(c['dbname']); print('DB OK', conn.info.server_version); conn.close()"
```
Expected: `BASE_DIR` prints the `InboundVerify` project root (the dir containing `config.yaml`); `DB OK <pg version>`.
> If 12b fails with ModuleNotFoundError for a site module, run `.venv/Scripts/python.exe -m pip install -e .` again (editable finder refresh) and retry. If 12d still shows a line, that import was missed — rewrite it per Step 4/8 rules.
- [ ] **Step 13: Commit (only after user confirms)**
```bash
git add -A inbound_verify
git commit -m "refactor: move flat modules into inbound_verify package (Tier 1, behavior-identical)
- relocate 12 root .py into inbound_verify/ (sites/, cli/ subpackages)
- rewrite all internal imports to package-qualified
- fix paths.py BASE_DIR to anchor at project root
- add main() entry wrappers; register console_scripts
- drop site_ prefix on site modules; keep state_store/expected_undelivered names
Co-Authored-By: Claude <noreply@anthropic.com>"
```
> Do NOT commit until the user says to. This commit is the **safety baseline**; the manual end-to-end test (spec §7 gate) runs against this state before any Tier 2/Tier 3 work.
---
## Task 3: Docs sync (README + CLAUDE.md)
**Files:**
- Modify: `README.md` (§二 directory tree, §三 env prep, §五 run)
- Modify: `CLAUDE.md` (常用命令 section)
**Interfaces:**
- Consumes: the completed package from Task 2 (docs must describe the real new layout/commands).
- Produces: documentation matching the new invocation model. No code impact.
- [ ] **Step 1: Update README §二 directory tree**
Replace the tree block (README lines ~3349) with the actual new layout:
```
InboundVerify/
├── pyproject.toml # 打包 + 依赖 + console_scripts
├── inbound_verify/ # 源码包
│ ├── paths.py runtime.py state_store.py expected_undelivered.py store.py
│ ├── sites/ shunxin / baishi / zto / yunda / anneng
│ └── cli/ router交互菜单/ serverFastAPI 服务)
├── config.example.yaml / config.yaml
├── schema.sql
├── requirements.txt # pyproject 的静态镜像
├── downloads/ output/ state/
└── docs/
```
- [ ] **Step 2: Update README §三 env prep — add editable install**
In the env-prep command block (README lines ~5871), after `pip install -r requirements.txt`, add:
```bash
# 4. 以可编辑模式安装本包(注册 inbound-verify 等命令)
pip install -e .
```
(renumber the subsequent `cp config.example.yaml config.yaml` step).
- [ ] **Step 3: Update README §五 run — new commands**
Replace `python main_router.py` with:
```bash
# 交互菜单(任选其一)
python -m inbound_verify.cli.router
# 或装包后inbound-verify
```
- [ ] **Step 4: Update CLAUDE.md 常用命令**
In the 常用命令 section, change every `.venv/Scripts/python.exe <module>.py` to the package form:
- `main_router.py``python -m inbound_verify.cli.router` (or `inbound-verify`)
- `server.py``python -m inbound_verify.cli.server` (or `inbound-verify-server`)
- `db_store.py createdb|init|ingest|all``python -m inbound_verify.store createdb|init|ingest|all` (or `inbound-verify-db ...`)
- `site_anneng.py expected|actual``python -m inbound_verify.sites.anneng expected|actual`
- `black`/`py_compile` targets → package paths (e.g. `-m black inbound_verify`)
Add a one-liner near the top of that section: `首次/拉取新代码后需 .venv/Scripts/python.exe -m pip install -e .`
- [ ] **Step 5: Verify docs render + commands are real**
```bash
grep -n "main_router.py\|python server.py\|python db_store.py\|site_anneng.py\|site_shunxin" README.md CLAUDE.md || echo "no stale old-path commands OK"
```
Expected: `no stale old-path commands OK` (every old invocation updated).
- [ ] **Step 6: Commit (only after user confirms)**
```bash
git add README.md CLAUDE.md
git commit -m "docs: update README and CLAUDE.md for package layout and console_scripts
Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Self-Review (completed)
- **Spec coverage:** spec §3 (layout) → Task 1+2; §4 (paths anchor) → Task 2 Step 2; §5 (import rules) → Task 2 Steps 39; §6 (entry/packaging) → Task 1 Step 2 + Task 2 Steps 6/8/9; §7 (verification gate) → Task 2 Step 12; §10 (docs) → Task 3. All covered.
- **Placeholder scan:** none — every step has exact code or an exact command with expected output. The 28 call-site rewrites are given as a deterministic prefix-drop rule + enumerated line numbers + verification grep (complete, not a placeholder).
- **Type/name consistency:** kept-module names (`state_store`, `expected_undelivered`, `runtime`, `paths`) used consistently across all import rewrites and call sites; renamed entries (`store`, `cli/router`, `cli/server`) consistent with pyproject `[project.scripts]`. `main()` signature consistent across router/server/store and console_scripts.

View File

@@ -0,0 +1,403 @@
# Tier 2: domain extract + rename + path dedup — Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Three targeted, behavior-preserving cleanups of the compare/config layer: (1) make `expected_undelivered` use `paths.DOWNLOAD_DIR/OUTPUT_DIR` instead of its own duplicate anchor; (2) extract shared site/file/column config into a new `domain` module; (3) rename `expected_undelivered``compare`.
**Architecture:** Pure refactor — move definitions, rewire imports, no logic change. Removes the duplicate path anchor that caused the Tier 1 hotfix bug (`17293be`), and the `db_store → expected_undelivered` coupling where the DB layer imported the whole compare engine just to read site config.
**Tech Stack:** Python ≥3.10, package `inbound_verify`, pandas, openpyxl, psycopg3.
## Global Constraints
- **Python ≥ 3.10**, package import name `inbound_verify`.
- **No behavior change** — pure refactor. Any logic change is a defect.
- **No test suite** (user decision). Per-task verification = `compileall` + **import smoke** (fresh process, no backend needed) + grep-for-stale-refs. NOT pytest. (The running backend is unaffected until a restart; a final end-to-end re-confirm is done once at the end of Tier 2.)
- **No auto-commit** (project rule). Each task's commit step runs only after the user says "提交". Commit messages in English.
- **Black-format** every changed `.py` (global rule).
- All changes inside the `InboundVerify` submodule; commits local on `dev`.
- Venv interpreter: `.venv/Scripts/python.exe` (absolute: `D:/projects/LogisticsHubIPA/InboundVerify/.venv/Scripts/python.exe`). Run commands from `D:/projects/LogisticsHubIPA/InboundVerify/`.
**Reference spec:** `docs/superpowers/specs/2026-07-23-package-restructure-design.md` §8 Tier 2.
**Out of scope (deferred):** `config.py` centralization (consolidating the 6 `open(CONFIG_PATH)+yaml.safe_load` sites in store/runtime×2/anneng/router). It is the churniest Tier 2 item (6 files) for the least marginal value — pure DRY, no pain addressed — and under no-tests more churn = more risk. Reconsider after the rest of Tier 2 lands.
---
## File Structure
```
inbound_verify/
├── domain.py # NEW (Task 2): shared site/file/column config — leaf module
├── paths.py # unchanged (already the single anchor)
├── expected_undelivered.py# Task 1: use paths.* ; Task 2: import config from domain ; Task 3: renamed → compare.py
├── compare.py # (after Task 3) was expected_undelivered.py — compare engine + report
├── store.py # Task 2: import config from domain (not eu); Task 3: import _read_business_dates from compare
├── runtime.py # Task 2: SITE_UNDELIVERED_FILE from domain; Task 3: write_site_file from compare
└── cli/router.py # Task 3: compare.main()
```
**`domain.py` responsibility:** pure data — `ALL_REPORT_SITES`, `SITE_UNDELIVERED_FILE`, `BAISHI_FILE`, `BAISHI_COLUMNS`, `arrived_pieces_zhongtong`, `arrived_pieces_by_cols`, `STATIONS`, `_site_cfg`. No `state_store` dependency, no file I/O. Leaf module.
---
## Task 1: Path dedup — expected_undelivered uses paths.DOWNLOAD_DIR/OUTPUT_DIR
**Files:**
- Modify: `inbound_verify/expected_undelivered.py` (lines 43-48 defs; usages at 146,147,298,338,374,400,669,675,676,690)
**Interfaces:**
- Consumes: `paths.DOWNLOAD_DIR`, `paths.OUTPUT_DIR` (already exist, anchored at project root).
- Produces: `expected_undelivered` no longer defines `BASE/DOWNLOADS/OUTPUT`; keeps `OUTFILE` (now `= join(OUTPUT_DIR, "应到未到数据.xlsx")`). The 6 `DOWNLOADS` and 1 `OUTPUT` usages point at the `paths.*` constants — same resolved values as the current hotfixed anchor, so behavior identical.
- [ ] **Step 1: Add the paths import to the top import block**
In `inbound_verify/expected_undelivered.py`, after the existing `import os` (top of file), add a line importing the two path constants. (The file currently has no `paths` import — it used its own BASE.)
Add:
```python
from inbound_verify.paths import DOWNLOAD_DIR, OUTPUT_DIR
```
(place it alongside the other imports near the top, e.g. right after `import os`.)
- [ ] **Step 2: Replace the self-anchored BASE/DOWNLOADS/OUTPUT/OUTFILE block**
Replace this block (currently lines ~43-48):
```python
# 包内文件:上两级 = 项目根(与 paths.BASE_DIR 一致;站点下载落在 <root>/downloads
# 注:本模块自带锚点是 paths.py 的重复Tier 2 计划改为直接引用 paths.DOWNLOAD_DIR/OUTPUT_DIR。
BASE = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DOWNLOADS = os.path.join(BASE, "downloads")
OUTPUT = os.path.join(BASE, "output")
OUTFILE = os.path.join(OUTPUT, "应到未到数据.xlsx")
```
with:
```python
# 比对报表输出文件(路径锚定统一走 paths.py
OUTFILE = os.path.join(OUTPUT_DIR, "应到未到数据.xlsx")
```
- [ ] **Step 3: Replace all `DOWNLOADS` usages with `DOWNLOAD_DIR`**
Apply `DOWNLOADS``DOWNLOAD_DIR` (6 occurrences, at lines 146, 147, 298, 338, 669, 675, 676). Use replace-all on the token `DOWNLOADS`.
- [ ] **Step 4: Replace the remaining `OUTPUT` usage (the makedirs line)**
At line ~374, replace:
```python
os.makedirs(OUTPUT, exist_ok=True)
```
with:
```python
os.makedirs(OUTPUT_DIR, exist_ok=True)
```
(`OUTFILE` at lines 400/690 is unchanged — it's a different token, already redefined in Step 2.)
- [ ] **Step 5: Verify syntax + no stale BASE/DOWNLOADS/OUTPUT**
```bash
cd /d/projects/LogisticsHubIPA/InboundVerify
.venv/Scripts/python.exe -m py_compile inbound_verify/expected_undelivered.py
grep -nE "\b(BASE|DOWNLOADS|OUTPUT)\b" inbound_verify/expected_undelivered.py || echo "no stale BASE/DOWNLOADS/OUTPUT OK"
```
Expected: py_compile silent; grep prints `no stale BASE/DOWNLOADS/OUTPUT OK` (OUTFILE is a different token, won't match).
- [ ] **Step 6: Import smoke + paths equivalence**
```bash
.venv/Scripts/python.exe -c "from inbound_verify import expected_undelivered as eu; from inbound_verify.paths import DOWNLOAD_DIR, OUTPUT_DIR; import os; print('OUTFILE dir matches OUTPUT_DIR:', os.path.dirname(eu.OUTFILE)==OUTPUT_DIR)"
.venv/Scripts/python.exe -c "import inbound_verify.cli.router, inbound_verify.cli.server, inbound_verify.store, inbound_verify.runtime; print('import smoke OK')"
```
Expected: `OUTFILE dir matches OUTPUT_DIR: True` and `import smoke OK`.
- [ ] **Step 7: Black + commit (after user confirms)**
```bash
.venv/Scripts/python.exe -m black inbound_verify/expected_undelivered.py
git add inbound_verify/expected_undelivered.py
git commit -m "refactor: expected_undelivered uses paths.DOWNLOAD_DIR/OUTPUT_DIR (Tier 2)
Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 2: Extract `domain.py` (shared site/file/column config)
**Files:**
- Create: `inbound_verify/domain.py`
- Modify: `inbound_verify/expected_undelivered.py` (remove the moved block, import from domain)
- Modify: `inbound_verify/store.py` (config from domain, not eu)
- Modify: `inbound_verify/runtime.py` (SITE_UNDELIVERED_FILE from domain)
**Interfaces:**
- Consumes: nothing new (moves existing definitions verbatim).
- Produces: `inbound_verify.domain` exposing `ALL_REPORT_SITES`, `SITE_UNDELIVERED_FILE`, `BAISHI_FILE`, `BAISHI_COLUMNS`, `arrived_pieces_zhongtong(df)`, `arrived_pieces_by_cols(wb_col, piece_col)`, `STATIONS` (list of dicts), `_site_cfg(name)`. These are the exact same objects that lived in `expected_undelivered` lines 50-135.
- [ ] **Step 1: Create `inbound_verify/domain.py`**
Create `inbound_verify/domain.py` with this exact content (moved verbatim from expected_undelivered.py lines 50-135, plus the `defaultdict` import it needs):
```python
# -*- coding: utf-8 -*-
"""domain.py — 站点 / 文件名 / 列映射的共享配置(单一来源)。
比对compare与入库store都依赖这套配置抽出独立 leaf 模块,
让 store 不必为读配置而依赖整个比对引擎。纯数据,无 state_store / 文件 IO 依赖。
"""
from collections import defaultdict
# 汇总报表覆盖的全部站点4 站在前、百世在末;汇总页图表只取 4 站)
ALL_REPORT_SITES = ["顺心", "中通", "韵达", "安能", "百世"]
# 4 站单站未到明细文件名(百世未到文件由站点直接产出,名为 BAISHI_FILE
SITE_UNDELIVERED_FILE = "{name}-未到数据.xlsx"
BAISHI_FILE = "百世-应到未到货物数据.xlsx"
BAISHI_COLUMNS = ["类型", "子单号", "运单号", "最新扫描记录"]
def arrived_pieces_zhongtong(df):
"""中通实到「运单号」为复合串H + 运单号(12) + 总数(4) + 顺序(4))。
基号 = v[:-8](与应到表运单号对齐),单件 = 整串(每串即一件)。"""
res = defaultdict(set)
for v in df["运单号"]:
v = str(v).strip()
if len(v) > 8 and v[-4:].isdigit():
res[v[:-8]].add(v) # 以完整复合串作为“已到单号”存入
return res
def arrived_pieces_by_cols(wb_col, piece_col):
"""顺心 / 韵达 / 安能:按干净运单列分组,单件 = 子单号 / 扫描单号。
wb_col实到表中与应到运单号对齐的干净列
(顺心=运单号 / 韵达=主单号 / 安能=所属单号)
piece_col实到表中每件货物的单号列子单号 / 扫描单号)"""
def parse(df):
res = defaultdict(set)
for m, s in zip(df[wb_col], df[piece_col]):
m, s = str(m).strip(), str(s).strip()
if m and s:
res[m].add(s)
return res
return parse
STATIONS = [
{
"name": "中通",
"exp": "中通-应到货物数据.xlsx",
"act": "中通-实到货物数据.xlsx",
"exp_qty": "交接件数", # 应到件数口径:交接件数(非录单件数)
"exp_wb": "运单号", # 应到表运单号列(兼作去重键)
"exp_jd": "交接单号", # 未到数据需展示的交接单号
"arrived_pieces": arrived_pieces_zhongtong,
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "顺心",
"exp": "顺心-应到货物数据.xlsx",
"act": "顺心-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("运单号", "子单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "韵达",
"exp": "韵达-应到货物数据.xlsx",
"act": "韵达-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("主单号", "子单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "安能",
"exp": "安能-应到货物数据.xlsx",
"act": "安能-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("所属单号", "扫描单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
]
def _site_cfg(name):
"""按名称取 4 站配置(百世不在 STATIONS返回 None"""
return next((c for c in STATIONS if c["name"] == name), None)
```
- [ ] **Step 2: Remove the moved block from expected_undelivered.py and import from domain**
In `inbound_verify/expected_undelivered.py`:
- Add to the top import block:
```python
from inbound_verify.domain import (
ALL_REPORT_SITES,
BAISHI_COLUMNS,
BAISHI_FILE,
SITE_UNDELIVERED_FILE,
STATIONS,
_site_cfg,
arrived_pieces_by_cols,
arrived_pieces_zhongtong,
)
```
- Delete the now-duplicated definitions (the block from `ALL_REPORT_SITES = ...` through the end of `_site_cfg`, i.e. old lines ~50-135 — the comment header `# 汇总报表...` through `return next(...)`). These now live in domain.py. The `from collections import defaultdict` import in expected_undelivered.py can stay (harmless) or be removed if unused — check with grep after.
- [ ] **Step 3: store.py — take config from domain, _read_business_dates still from eu**
In `inbound_verify/store.py`:
- Add to imports:
```python
from inbound_verify.domain import BAISHI_FILE, _site_cfg
```
- Replace `cfg = eu._site_cfg(site)` (2 occurrences, lines 264 and 296) → `cfg = _site_cfg(site)`.
- Replace `eu.BAISHI_FILE` (2 occurrences, lines 335 and 337) → `BAISHI_FILE`.
- Leave `eu._read_business_dates(...)` (line 213) unchanged — that's compare behavior, stays accessed via the compare module (`eu`). The `from inbound_verify import expected_undelivered as eu` import stays for this.
- [ ] **Step 4: runtime.py — SITE_UNDELIVERED_FILE from domain**
In `inbound_verify/runtime.py`:
- Add to imports:
```python
from inbound_verify.domain import SITE_UNDELIVERED_FILE
```
- Replace (line ~448):
```python
DOWNLOAD_DIR, expected_undelivered.SITE_UNDELIVERED_FILE.format(name=site)
```
with:
```python
DOWNLOAD_DIR, SITE_UNDELIVERED_FILE.format(name=site)
```
- Leave `expected_undelivered.write_site_file(site)` (line ~446) and `expected_undelivered.main()` (line ~476) unchanged (compare behavior).
- [ ] **Step 5: Verify — compileall + import smoke + domain is leaf**
```bash
cd /d/projects/LogisticsHubIPA/InboundVerify
.venv/Scripts/python.exe -m compileall -q inbound_verify
.venv/Scripts/python.exe -c "import inbound_verify.domain as d; print('STATIONS:', [s['name'] for s in d.STATIONS]); print('_site_cfg(中通):', d._site_cfg('中通')['exp']); print('BAISHI_FILE:', d.BAISHI_FILE)"
.venv/Scripts/python.exe -c "import inbound_verify.cli.router, inbound_verify.cli.server, inbound_verify.store, inbound_verify.runtime, inbound_verify.expected_undelivered; print('import smoke OK')"
```
Expected: STATIONS lists `['中通', '顺心', '韵达', '安能']`; `_site_cfg('中通')['exp']` = `中通-应到货物数据.xlsx`; `BAISHI_FILE` = `百世-应到未到货物数据.xlsx`; `import smoke OK`.
- [ ] **Step 6: Black + commit (after user confirms)**
```bash
.venv/Scripts/python.exe -m black inbound_verify/domain.py inbound_verify/expected_undelivered.py inbound_verify/store.py inbound_verify/runtime.py
git add inbound_verify/domain.py inbound_verify/expected_undelivered.py inbound_verify/store.py inbound_verify/runtime.py
git commit -m "refactor: extract domain.py (shared site/file/colmap config) from expected_undelivered
Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Task 3: Rename `expected_undelivered.py` → `compare.py`
**Files:**
- Move: `inbound_verify/expected_undelivered.py``inbound_verify/compare.py` (git mv)
- Modify: `inbound_verify/store.py`, `inbound_verify/runtime.py`, `inbound_verify/cli/router.py` (import + call-site rewrites)
**Interfaces:**
- Consumes: Task 2's `domain` (compare still uses it).
- Produces: module is `inbound_verify.compare`; all public names (`main`, `write_site_file`, `_read_business_dates`) unchanged. `expected_undelivered` no longer exists as a module name.
- [ ] **Step 1: git mv the module (history preserved)**
```bash
cd /d/projects/LogisticsHubIPA/InboundVerify
git mv inbound_verify/expected_undelivered.py inbound_verify/compare.py
```
- [ ] **Step 2: store.py — import compare instead of expected_undelivered**
In `inbound_verify/store.py`, replace:
```python
from inbound_verify import expected_undelivered as eu
```
with:
```python
from inbound_verify import compare
```
and replace the one call site (line ~213):
```python
return eu._read_business_dates(ALL_SITES + ["百世"]) or {}
```
with:
```python
return compare._read_business_dates(ALL_SITES + ["百世"]) or {}
```
(`_site_cfg` and `BAISHI_FILE` already come from `domain` after Task 2 — no change there.)
- [ ] **Step 3: runtime.py — import compare, fix write_site_file/main**
In `inbound_verify/runtime.py`, replace:
```python
from inbound_verify import expected_undelivered
```
with:
```python
from inbound_verify import compare
```
Replace `expected_undelivered.write_site_file(site)` (line ~446) → `compare.write_site_file(site)`.
Replace `(expected_undelivered.main() or True)` (line ~476) → `(compare.main() or True)`.
- [ ] **Step 4: cli/router.py — import compare, fix main()**
In `inbound_verify/cli/router.py`, replace:
```python
from inbound_verify import expected_undelivered
```
with:
```python
from inbound_verify import compare
```
Replace `expected_undelivered.main()` (line ~34, inside `run_undelivered_compare`) → `compare.main()`.
- [ ] **Step 5: Verify — no stale references + import smoke**
```bash
cd /d/projects/LogisticsHubIPA/InboundVerify
.venv/Scripts/python.exe -m compileall -q inbound_verify
grep -rn "expected_undelivered" inbound_verify || echo "no stale expected_undelivered refs OK"
.venv/Scripts/python.exe -c "import inbound_verify.compare, inbound_verify.cli.router, inbound_verify.cli.server, inbound_verify.store, inbound_verify.runtime; print('import smoke OK')"
```
Expected: compileall silent; grep prints `no stale expected_undelivered refs OK`; `import smoke OK`.
- [ ] **Step 6: Black + commit (after user confirms)**
```bash
.venv/Scripts/python.exe -m black inbound_verify
git add -A inbound_verify
git commit -m "refactor: rename expected_undelivered to compare
Co-Authored-By: Claude <noreply@anthropic.com>"
```
---
## Final end-to-end re-confirm (once, after all 3 tasks; requires backend restart + re-login)
After Task 3 commits, the running backend still has the old modules loaded. To confirm end-to-end behavior is unchanged on the refactored code:
- [ ] **Restart backend** (TaskStop current → clean ms-playwright/anneng orphans → `python -m inbound_verify.cli.server`), re-login all sites.
- [ ] **Re-run the gate from the Tier 1 manual test:** trigger `undelivered` for 顺心/中通/韵达/安能 via API → expect 4× success + 4× `*-未到数据.xlsx`; trigger `__compare__` → expect `output/应到未到数据.xlsx`; run `python -m inbound_verify.store ingest` → expect rows UPSERTed. All green = Tier 2 behavior-identical, done.
> If you want to skip the re-login cost: the per-task import-smoke gates already prove the import graph is correct and the changes are behavior-preserving moves. The e2e re-confirm is belt-and-suspenders.
---
## Self-Review (completed)
- **Spec coverage:** spec §8 Tier 2 — path dedup → Task 1; domain extract → Task 2; rename → compare → Task 3. config.py explicitly deferred (noted with rationale). All spec items addressed or consciously deferred.
- **Placeholder scan:** none — every step has exact code or exact commands with expected output. The domain.py content is the verbatim extracted block.
- **Type/name consistency:** `_site_cfg`, `BAISHI_FILE`, `SITE_UNDELIVERED_FILE`, `STATIONS`, `write_site_file`, `main`, `_read_business_dates` referenced consistently across tasks. Task 2 routes config to `domain` and leaves behavior (`_read_business_dates`, `write_site_file`, `main`) in compare — verified against store.py/runtime.py/router.py usages. Task 3 renames the module but preserves all public names.

View File

@@ -0,0 +1,522 @@
# 下载后自动入库钩子 Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** 把 PostgreSQL 入库动作挂到 `runtime.dispatch_task` 下载成功分支使每次下载成功后自动、尽力而为、kind 级地把刚下载的数据 UPSERT 进 PG。
**Architecture:** 新增 `store.ingest_task(site, kind)`kind 级路由,复用现有 `_ingest_*`+ `runtime._persist_to_db(site, kind)`(同步内联钩子,懒导入 store绝不外抛+ `state_store.ingest_state` 表(结果可查,经 `/status` 暴露)+ 配置开关/超时。详见 spec `docs/superpowers/specs/2026-07-24-ingest-hook-design.md`
**Tech Stack:** Python 3.10+、psycopg(v3)、sqlite3、Playwright(不动)、FastAPI(仅 `/status` 加字段)、pyyaml。包以可编辑模式安装`pip install -e .`),命令走 `.venv/Scripts/python.exe -m inbound_verify...`
## Global Constraints
- **Python 一律走项目虚拟环境**`.venv/Scripts/python.exe`(全局 CLAUDE.md 协议)。命令默认在 `InboundVerify/` 根目录执行。
- **无 pytest 测试套件**(本仓库约定,覆盖 writing-plans 默认的 TDD-pytest 步骤):每个任务的验证用 `py_compile` + `black` + 导入冒烟 + 功能/手动校验,不写 pytest。
- **改完任何 .py 必须跑 Black**`.venv/Scripts/python.exe -m black inbound_verify`
- **不自动提交/推送**(本仓库约定,覆盖 writing-plans 默认的"每任务即提交"):每个任务的 Commit 步骤**仅在用户明确说"提交"/"commit"时执行**;否则完成任务后停在待提交态,告知用户。
- **commit message 一律英文**(全局 CLAUDE.md
- **不改下载流程/比对逻辑**4 站 `-未到数据.xlsx` 不入库(既有设计);不上连接池/后台线程/补入重试。
- 导入约束:`runtime` 已 import `compare``store` 也 import `compare`;钩子里对 `store` **懒导入**以回避成环。
---
## File Structure
| 文件 | 责任 | 本计划改动 |
|---|---|---|
| `inbound_verify/store.py` | PG 持久化(建库/建表/入库 CLI | 改 `_load_pg_config`/`_connect`;加 `ingest_enabled()``ingest_task()`CLI 加 `ingest-one` |
| `inbound_verify/state_store.py` | SQLite 状态持久化 | `init_db``ingest_state` 表;加 `set_ingest_state()``get_all_ingest_state()` |
| `inbound_verify/runtime.py` | 两种模式共享核心(含 dispatch_task | 加 `_persist_to_db()`;在 `dispatch_task` 成功分支调用 |
| `inbound_verify/cli/server.py` | FastAPI 服务模式 | `/status` 返回加 `ingest` 字段 |
| `config.example.yaml` | 配置模板(提交) | `postgres` 段加 `auto_ingest`/`connect_timeout_seconds`;修过期 `db_store.py``store.py` 注释 |
| `config.yaml` | 真实配置gitignored | 同上两键 + 修注释 |
| `README.md` | 用户文档 | DB CLI 命令列表加 `ingest-one` |
任务依赖Task 2 依赖 Task 1Task 4 依赖 Task 1+2+3Task 5 依赖 Task 3。Task 3 独立。建议顺序 1→2→3→4→5。
---
### Task 1: store.py 配置键 + 连接超时层
**Files:**
- Modify: `inbound_verify/store.py:48-78``_load_pg_config``_connect`
- Modify: `config.example.yaml:72-89``config.yaml`postgres 段)
**Interfaces:**
- Consumes: 无(配置层根基)
- Produces: `_load_pg_config()` 返回新增 `auto_ingest: bool``connect_timeout_seconds: int``_connect(dbname)` 连接带 `connect_timeout` 且会话级 `statement_timeout=30s`;新函数 `ingest_enabled() -> bool`。后续任务依赖 `ingest_enabled()` 与超时连接。
- [ ] **Step 1: 改 `_load_pg_config` 加两键**
`inbound_verify/store.py``_load_pg_config` 返回 dict`schema` 之后追加两键:
```python
return {
"host": pg.get("host", "127.0.0.1"),
"port": int(pg.get("port", 5432)),
"user": pg.get("user", "postgres"),
"password": pg.get("password", ""),
"dbname": pg.get("dbname", "CQHXDB"),
"schema": pg.get("schema", "inbound_verify"),
"auto_ingest": bool(pg.get("auto_ingest", True)),
"connect_timeout_seconds": int(pg.get("connect_timeout_seconds", 5)),
}
```
- [ ] **Step 2: 改 `_connect` 加 connect_timeout + statement_timeout**
`_connect` 替换为(用 `options` 一次性设 search_path + statement_timeout等价于 spec 的 SET 但不引入额外事务):
```python
def _connect(dbname):
"""用关键字参数连接(避开 conninfo 对密码特殊字符的解析)。
options 设 search_path 到专用 schema + 会话级 statement_timeout=30s
cpolar 隧道上防失控查询connect_timeout 守连接阶段)。"""
c = _load_pg_config()
return psycopg.connect(
host=c["host"],
port=c["port"],
dbname=dbname,
user=c["user"],
password=c["password"],
options=f"-c search_path={c['schema']} -c statement_timeout=30s",
connect_timeout=c["connect_timeout_seconds"],
)
```
- [ ] **Step 3: 加 `ingest_enabled()` 薄封装**
`_connect` 之后、`# 建库 / 建表` 分节注释之前插入:
```python
def ingest_enabled():
"""是否启用下载后自动入库config.yaml postgres.auto_ingest默认 True
供 runtime 钩子判定开关,避免它伸手进 _load_pg_config。"""
return _load_pg_config()["auto_ingest"]
```
- [ ] **Step 4: 更新 store.py 模块 docstring 的命令列表**
把文件顶部 docstring 的命令行小节,在 `ingest` 行后补一行 `ingest-one`Task 2 会实现该命令docstring 先行):
```
python -m inbound_verify.store ingest [site] 入库全站或单站(幂等 UPSERT
python -m inbound_verify.store ingest-one <site> <kind> 仅入库指定站/类(钩子同款路由)
python -m inbound_verify.store all createdb → init → 全站 ingest 一条龙
```
- [ ] **Step 5: config.example.yaml 加两键 + 修过期注释**
`config.example.yaml` 的 postgres 段:把 `# PostgreSQL 数据持久化(到货核销数据入库,详见 db_store.py` 改为 `... 详见 store.py`;命令注释里的 `python db_store.py ...` 改为 `python -m inbound_verify.store ...`;在 `schema: inbound_verify` 之后追加:
```yaml
schema: inbound_verify
# 下载成功后自动入库(钩子,见 runtime._persist_to_dbfalse=跳过(无 PG/cpolar 的开发机)。
auto_ingest: true
# PG 连接超时cpolar 抖动时快速失败,不拖垮下载 worker。
connect_timeout_seconds: 5
```
- [ ] **Step 6: config.yaml 同步gitignored 本地文件)**
`config.yaml` 做同样两键追加,并把注释里的 `db_store.py` 改为 `store.py`。保留其真实凭据不动。
- [ ] **Step 7: 编译 + Black**
```
.venv/Scripts/python.exe -m py_compile inbound_verify/store.py
.venv/Scripts/python.exe -m black inbound_verify/store.py
```
Expected: py_compile 无输出black 报 `reformatted``left unchanged`
- [ ] **Step 8: 回归既有 CLI 不破**
```
.venv/Scripts/python.exe -c "from inbound_verify import store; c=store._load_pg_config(); assert c['auto_ingest'] is True and c['connect_timeout_seconds']==5; assert store.ingest_enabled() is True; print('config ok')"
```
Expected: `config ok`
- [ ] **Step 9: Commitgated**
仅当用户说"提交"时执行:
```bash
git add inbound_verify/store.py config.example.yaml
git commit -m "feat(store): add auto_ingest config + pg connect/statement timeouts"
```
`config.yaml` 已 gitignore不入提交。
---
### Task 2: store.ingest_task + ingest-one CLI
**Files:**
- Modify: `inbound_verify/store.py`(加 `ingest_task``main()``ingest-one` 分支)
**Interfaces:**
- Consumes: Task 1 的 `_connect`(超时)、`_load_pg_config`、既有 `_ingest_expected/_ingest_actual/_ingest_undelivered_baishi``_read_business_dates``ALL_SITES``_site_cfg`
- Produces: `ingest_task(site: str, kind: str) -> int`(返回入库总条数;`__compare__`/不支持组合返回 0
- [ ] **Step 1: 加 `ingest_task`**
`ingest(site)` 函数之后、`# 命令行` 分节之前插入:
```python
def ingest_task(site, kind):
"""按 (site, kind) 入库本次刚下载的文件(幂等 UPSERT返回总条数。
与 ingest(site) 的区别:只入本次刷新的那一类,避免重读写另一类文件(同步钩子里减少阻塞)。
kind 路由:
expected/actual 各入其列;
undelivered 百世 入未到;
undelivered 4 站 _site_undelivered_handler 内部连带下了 expected+actual故入两者
__compare__ / 其它组合 返回 0。
"""
if site == "__compare__":
return 0
dates = _read_business_dates()
total = 0
with _connect(_load_pg_config()["dbname"]) as conn:
with conn.cursor() as cur:
if kind == "expected":
total += _ingest_expected(cur, site, dates.get(site))
elif kind == "actual":
total += _ingest_actual(cur, site)
elif kind == "undelivered":
if site == "百世":
total += _ingest_undelivered_baishi(cur)
else: # 顺心/中通/韵达/安能
total += _ingest_expected(cur, site, dates.get(site))
total += _ingest_actual(cur, site)
# 其它组合(如 百世/expected正常不经钩子触发防御性返回 0
conn.commit()
return total
```
- [ ] **Step 2: `main()` 加 `ingest-one` 分支**
`main()``elif cmd == "all":` 分支之后、`else:` 之前插入:
```python
elif cmd == "ingest-one":
kind = sys.argv[3] if len(sys.argv) > 3 else None
if not site or kind not in ("expected", "actual", "undelivered"):
print("用法: python -m inbound_verify.store ingest-one <site> <expected|actual|undelivered>")
sys.exit(1)
total = ingest_task(site, kind)
print(f">> [ingest-one] {site}/{kind} 入库 {total}")
```
- [ ] **Step 3: 编译 + Black**
```
.venv/Scripts/python.exe -m py_compile inbound_verify/store.py
.venv/Scripts/python.exe -m black inbound_verify/store.py
```
Expected: 无编译错误。
- [ ] **Step 4: 路由正确性(手动,需 PG 已 createdb+init 且 downloads/ 有数据)**
逐条验证 kind 级只入对应文件:
```
.venv/Scripts/python.exe -m inbound_verify.store ingest-one 韵达 expected
```
Expected: stdout 只出现 `[应到] 韵达N 条运单`**不**出现 `[实到] 韵达` 行;结尾 `[ingest-one] 韵达/expected 入库 N 条`
```
.venv/Scripts/python.exe -m inbound_verify.store ingest-one 顺心 undelivered
```
Expected: stdout 同时出现 `[应到] 顺心``[实到] 顺心`undelivered→两者
```
.venv/Scripts/python.exe -m inbound_verify.store ingest-one 百世 undelivered
```
Expected: stdout 出现 `[未到] 百世`
> 若某站 downloads/ 无文件:`_ingest_*` 打印 `[跳过] ... 文件不存在`ingest_task 返回 0属正常非错误
- [ ] **Step 5: 无 PG 时的降级(手动,可选)**
临时把 `config.yaml``host` 改成不可达地址,重跑 Step 4 任一命令:应在 `connect_timeout_seconds`(默认 5s内报 psycopg 连接错误并退出(非 hang。验证后改回真实 host。
- [ ] **Step 6: Commitgated**
```bash
git add inbound_verify/store.py
git commit -m "feat(store): add kind-level ingest_task + ingest-one CLI subcommand"
```
---
### Task 3: state_store ingest_state 表 + 读写函数
**Files:**
- Modify: `inbound_verify/state_store.py:32-113``init_db` 加表)、`:219-247`(加两函数)
**Interfaces:**
- Consumes: 无(独立叶子,仅依赖 paths + sqlite3 + datetime
- Produces: `set_ingest_state(site, kind, ok, count=0, error=None) -> None``get_all_ingest_state() -> dict`(形状 `{site: {kind: {ok, ingested_at, count, error}}}`。Task 4 与 Task 5 依赖这两个。
- [ ] **Step 1: `init_db` 加 `ingest_state` 表**
`init_db` 内、`site_settings` 表 CREATE 之后(`conn.commit()` 之前)插入:
```python
conn.execute("""
CREATE TABLE IF NOT EXISTS ingest_state (
site TEXT,
kind TEXT,
ok INTEGER,
ingested_at TEXT,
count INTEGER,
error TEXT,
PRIMARY KEY (site, kind)
)
""")
```
- [ ] **Step 2: 加 `set_ingest_state` 与 `get_all_ingest_state`**
`get_all_status()` 函数之后插入:
```python
def set_ingest_state(site, kind, ok, count=0, error=None):
"""记录一次入库结果UPSERT。ok: boolcount: 入库条数error: 失败原因或 None。"""
with sqlite3.connect(STATE_DB_PATH) as conn:
conn.execute(
"INSERT INTO ingest_state (site, kind, ok, ingested_at, count, error) "
"VALUES (?, ?, ?, ?, ?, ?) "
"ON CONFLICT(site, kind) DO UPDATE SET "
"ok=excluded.ok, ingested_at=excluded.ingested_at, "
"count=excluded.count, error=excluded.error",
(site, kind, 1 if ok else 0, _now(), int(count or 0), error or ""),
)
conn.commit()
def get_all_ingest_state():
"""返回 {site: {kind: {ok, ingested_at, count, error}}};库不存在返回 {}"""
if not os.path.exists(STATE_DB_PATH):
return {}
with sqlite3.connect(STATE_DB_PATH) as conn:
rows = conn.execute(
"SELECT site, kind, ok, ingested_at, count, error FROM ingest_state"
).fetchall()
out = {}
for site, kind, ok, ingested_at, count, error in rows:
out.setdefault(site, {})[kind] = {
"ok": bool(ok),
"ingested_at": ingested_at or "",
"count": int(count or 0),
"error": error or "",
}
return out
```
- [ ] **Step 3: 编译 + Black**
```
.venv/Scripts/python.exe -m py_compile inbound_verify/state_store.py
.venv/Scripts/python.exe -m black inbound_verify/state_store.py
```
Expected: 无错误。
- [ ] **Step 4: set/get 往返冒烟**
```
.venv/Scripts/python.exe -c "from inbound_verify import state_store as s; s.init_db(); s.set_ingest_state('韵达','expected',True,count=42); s.set_ingest_state('韵达','expected',False,error='boom'); d=s.get_all_ingest_state(); r=d['韵达']['expected']; assert r['ok'] is False and r['count']==42 and r['error']=='boom' and r['ingested_at']; print('ingest_state ok')"
```
Expected: `ingest_state ok`(验证 UPSERT 覆盖:第二次写把 ok 改 Falsecount 保留 42error 写入)。
- [ ] **Step 5: Commitgated**
```bash
git add inbound_verify/state_store.py
git commit -m "feat(state_store): add ingest_state table + set/get helpers"
```
---
### Task 4: runtime._persist_to_db 钩子 + 挂到 dispatch_task
**Files:**
- Modify: `inbound_verify/runtime.py`(加 `_persist_to_db`,位置在 `_record_business_date` 之后、`dispatch_task` 之前;改 `dispatch_task` 成功分支 `:552-557`
**Interfaces:**
- Consumes: Task 1 `store.ingest_enabled()`、Task 2 `store.ingest_task(site, kind)`、Task 3 `state_store.set_ingest_state(...)`;既有 `state_store`runtime 已 import
- Produces: `_persist_to_db(site, kind)``dispatch_task` 在下载成功后调用;下载任务的成功判定**不变**。
- [ ] **Step 1: 加 `_persist_to_db`**
`_record_business_date` 函数之后(`def dispatch_task` 之前)插入:
```python
def _persist_to_db(site, kind):
"""下载成功后把本次数据入库 PostgreSQL尽力而为绝不外抛不影响任务判定
- __compare__ 无源数据,跳过。
- auto_ingest=false 时跳过(无 PG/cpolar 的开发机)。
- 懒导入 store 以回避 import 顺序store↔compare 与 runtime↔compare 共存)。
- 结果写 state_store.ingest_state供 /api/status 反映入库健康。
所有写库/写状态都包 try/except失败仅告警绝不改变 dispatch_task 的 SUCCESS 判定。"""
if site == "__compare__":
return
try:
from inbound_verify import store # 懒导入:冷路径(每下载一次),回避成环
except Exception as e:
print(f">> [warn] 入库模块不可用: {e}")
return
if not store.ingest_enabled():
print(">> [入库] 已关闭 (auto_ingest=false),跳过")
return
try:
count = store.ingest_task(site, kind)
state_store.set_ingest_state(site, kind, ok=True, count=count)
print(f">> [入库] {site}/{kind} 成功,{count}")
except Exception as e:
print(f">> [warn] 入库失败 {site}/{kind}: {e}")
try:
state_store.set_ingest_state(site, kind, ok=False, error=str(e))
except Exception as e2:
print(f">> [warn] 写入库状态也失败: {e2}")
```
- [ ] **Step 2: 挂到 `dispatch_task` 成功分支**
`dispatch_task` 内的成功分支:
```python
_record_business_date(site, kind)
return (state_store.TASK_SUCCESS, None)
```
改为:
```python
_record_business_date(site, kind)
_persist_to_db(site, kind)
return (state_store.TASK_SUCCESS, None)
```
`_persist_to_db` 绝不外抛,故不会被外层 `except Exception` 误判为任务失败。)
- [ ] **Step 3: 编译 + Black + 导入冒烟**
```
.venv/Scripts/python.exe -m py_compile inbound_verify/runtime.py
.venv/Scripts/python.exe -m black inbound_verify/runtime.py
.venv/Scripts/python.exe -c "from inbound_verify.runtime import dispatch_task, _persist_to_db; print('runtime import ok')"
```
Expected: `runtime import ok`
- [ ] **Step 4: 端到端(手动,需站点已登录)**
任选一种模式触发一次真实下载并观察钩子:
- 服务模式:`POST /tasks` `{"site":"韵达","kind":"expected"}`(或经 dashboard 触发),下载完成后看 worker stdout
- 成功:`>> [入库] 韵达/expected 成功N 条`
- 失败(如 PG 未 init`>> [warn] 入库失败 ...`,且任务本身仍 `success``GET /tasks/{id}` 验证)。
- 交互模式:菜单 `[6]` 韵达应到,完成后看同样的 `[入库]` 行。
并查 `state/state.db`
```
.venv/Scripts/python.exe -c "from inbound_verify import state_store as s; print(s.get_all_ingest_state())"
```
Expected: 含 `韵达`/`expected` 的记录,`ok` 与 stdout 一致。
- [ ] **Step 5: 关开关降级(手动)**
`config.yaml``auto_ingest``false`再触发一次下载stdout 应出现 `>> [入库] 已关闭 (auto_ingest=false),跳过`,且不连 PG任务仍 SUCCESS。验证后改回 `true`
- [ ] **Step 6: Commitgated**
```bash
git add inbound_verify/runtime.py
git commit -m "feat(runtime): auto-ingest hook after successful download"
```
---
### Task 5: /status 暴露 ingest 态 + README
**Files:**
- Modify: `inbound_verify/cli/server.py:189-196``get_status`
- Modify: `README.md`DB CLI 命令列表)
**Interfaces:**
- Consumes: Task 3 `state_store.get_all_ingest_state()`
- Produces: `GET /status` 返回体新增 `ingest` 字段。
- [ ] **Step 1: `/status` 加 `ingest` 字段**
`get_status` 的返回 dict
```python
return {
"worker_ready": worker_state["ready"],
"worker_error": worker_state["error"],
"sites": state_store.get_all_status(),
}
```
改为:
```python
return {
"worker_ready": worker_state["ready"],
"worker_error": worker_state["error"],
"sites": state_store.get_all_status(),
"ingest": state_store.get_all_ingest_state(),
}
```
并把 docstring 顺手补一句(可选):`"""各站登录态 + 数据态 + 入库态(前端状态盘用),另含 worker 就绪状态。"""`
- [ ] **Step 2: README DB CLI 命令列表加 `ingest-one`**
`README.md` 第五章"运行"下、DB CLI 注释行(`# 或inbound-verify-db createdb|init|ingest|all` 附近)补一行说明自动入库 + 新命令:
```
# DB CLI建库 / 初始化 / 灌数据 / 全流程 / 单站单类
.venv/Scripts/python.exe -m inbound_verify.store createdb # 或 init | ingest | ingest-one <site> <kind> | all
# 注下载成功后会自动入库postgres.auto_ingest默认开ingest-one 用于手动重灌指定站/类。
```
- [ ] **Step 3: 编译 + Black**
```
.venv/Scripts/python.exe -m py_compile inbound_verify/cli/server.py
.venv/Scripts/python.exe -m black inbound_verify/cli/server.py
```
Expected: 无错误。
- [ ] **Step 4: `/status` 含 ingest 字段(手动,服务模式已启动)**
```
curl -s http://127.0.0.1:8000/status | python -m json.tool
```
(或浏览器 `http://127.0.0.1:8000/docs``/status`Expected: 返回体含 `"ingest": {...}` 键(无入库记录时为 `{}`Task 4 跑过后会有值)。经 dashboard 的 `GET /api/status`(代理)同样可见。
- [ ] **Step 5: 全量回归冒烟**
```
.venv/Scripts/python.exe -m py_compile inbound_verify
.venv/Scripts/python.exe -m black inbound_verify
.venv/Scripts/python.exe -c "import inbound_verify.store, inbound_verify.state_store, inbound_verify.runtime, inbound_verify.cli.server; print('all imports ok')"
```
Expected: `all imports ok`black 无 diff。
- [ ] **Step 6: Commitgated**
```bash
git add inbound_verify/cli/server.py README.md
git commit -m "feat(server): expose ingest state in /status; doc auto-ingest in README"
```
---
## Self-Review 结论
- **Spec 覆盖**spec §4.1 → Task 1+2§4.2 → Task 4§4.3 → Task 3§4.4 → Task 5§4.5 配置 → Task 1config 两文件§7 测试手段 → 各任务手动步ingest-one 路由、端到端、降级、回归)。无遗漏。
- **占位符**:无 TBD/TODO每步含完整代码与确切命令。
- **类型/命名一致**`ingest_task(site, kind)``ingest_enabled()``set_ingest_state(site, kind, ok, count=0, error=None)``get_all_ingest_state()``_persist_to_db(site, kind)` 在各任务间签名一致;`/status` 字段名 `ingest``get_all_ingest_state` 返回一致。
- **已标注偏离**(a) 无 pytest → 用 py_compile/black/冒烟/手动代替Global Constraints(b) 不自动提交 → Commit 步骤 gatedGlobal Constraints + 每任务注明);(c) `_connect``options` 设 statement_timeout 取代 spec 的 `SET`等价、无额外事务Task 1 Step 2 注释说明)。

View File

@@ -0,0 +1,274 @@
# InboundVerify 包化重构设计
- **日期**:2026-07-23
- **方案**:C(规范)—— `pyproject.toml` + `[project.scripts]` + 根级包 `inbound_verify/`,根目录不留 `.py` 薄壳
- **力度**:彻底(建包 + 定点小改进 + 大文件拆分),但**大文件拆分(Tier 3)前置手动测试闸门**
- **状态**:已与用户对齐,待 spec 评审
---
## 1. 背景与目标
当前 13 个 Python 模块(共约 6500 行)全部平铺在仓库根目录,随脚本量增长结构混乱。本设计将其重组为一个规范的、可 `pip install -e .` 安装的 Python 包。
**目标**
1. 扁平脚本 → `inbound_verify/` 包(`sites/``cli/` 两个子包,其余平铺包根,避免一层只放一两个文件的过度嵌套)。
2. 全部内部 import 改为包内绝对引用。
3. `pyproject.toml` 打包,`[project.scripts]` 暴露命令入口;**根目录不留 `.py` 薄壳**(规范要求)。
4. 顺手抽离共享配置、消重(定点小改进,Tier 2)。
5. 大文件拆分作为**最后、可选、风险隔离**的阶段(Tier 3),且必须先过手动测试闸门。
**已确认约束**
- InboundVerify 是 git **子模块**;改动在子模块内提交,父仓库 `LogisticsHubIPA` 仅跟踪子模块指针。
- **不加测试套件**(用户决定);验证靠 `compileall` + 导入冒烟 + grep 查残留引用 + 手动端到端测试。
- 入口走规范:无根薄壳;一次性 `pip install -e .` 后用命令或 `python -m` 启动。
- **不做 src/ 布局**(内部工具、不发 PyPI、无测试,边际价值有限)。
- 本仓库约定:**不自动提交 / 不自动推送**;改动等用户明确说"提交"再 commit/push(本约定覆盖 brainstorming 默认的"写完即提交")。
---
## 2. 现状分析
### 2.1 依赖分层(自底向上)
```
paths ← 万物之基(被所有模块 import)
state_store ← SQLite 状态持久化
expected_undelivered ← 离线比对(兼职存站点/文件名/列映射共享配置,被 db_store 复用 —— 耦合点)
site_*.py(×5) ← 各依赖 paths + state_store
runtime ← 编排核心(启动/派发/心跳,依赖上面全部)
main_router / server / db_store ← 三入口(均有 __main__/CLI)
```
### 2.2 关键发现
1. **`paths.py` 是地雷**:`BASE_DIR = dirname(abspath(__file__))`——paths.py 在哪,根就在哪。搬进子目录后必须改为上跳一级,否则 `downloads/``output/``config.yaml``state/state.db``schema.sql` 全部跑偏。
2. **import 全是扁平顶层**(`import site_shunxin``from paths import …`),且**没有任何地方按字符串名引用模块**(`TASK_HANDLERS` 用函数引用、`dispatch_task` 用 site/kind 字典),改写纯机械。
3. **重复代码**:`with_retry``_remove_if_exists` 在 5 个站点逐字重复(CLAUDE.md 注明"现阶段刻意不优化结构")。
4. **路径常量重复**:`expected_undelivered.py` 自带 `BASE/DOWNLOADS/OUTPUT`,与 `paths.py` 重复。
5. **耦合**:`db_store` 为读站点配置而 `import expected_undelivered`(为了配置而依赖整个比对引擎)。
---
## 3. 目标目录结构
```
InboundVerify/
├── pyproject.toml # 新增:打包 + 依赖 + console_scripts
├── config.example.yaml
├── config.yaml # 留根(gitignored)
├── schema.sql # 留根(store 按绝对路径读)
├── requirements.txt # 保留为静态镜像(pyproject 为准)
├── README.md / CLAUDE.md / .gitignore
├── downloads/ output/ state/ # 运行时数据,留根
├── docs/
└── inbound_verify/ # ← 包
├── __init__.py
├── __main__.py # 可选:python -m inbound_verify → 交互菜单
├── paths.py # 锚点改为指向项目根(§4)
├── config.py # 新增(Tier2):集中 load_config()
├── domain.py # 新增(Tier2):从 expected_undelivered 抽出的共享站点/文件/列映射
├── runtime.py # ← runtime.py(编排核心,整体保留不拆)
├── state_store.py # ← state_store.py(保留名,仅搬运)
├── expected_undelivered.py # ← 搬运;Tier2 改名 compare.py + 抽 domain.py
├── store.py # ← db_store.py(叶子,改名无 churn)
├── sites/
│ ├── __init__.py
│ ├── shunxin.py baishi.py zto.py yunda.py # ← site_*.py(去 site_ 前缀)
│ └── anneng.py # ← site_anneng.py(Tier3 可选再拆成子包)
└── cli/
├── __init__.py
├── router.py # ← main_router.py(叶子,改名无 churn)
└── server.py # ← server.py(叶子)
```
### 3.1 搬运映射表(全部 `git mv` 保历史)
| 现在 | Tier 1 后 | 调用点改动 |
|---|---|---|
| `paths.py` | `inbound_verify/paths.py` | 无(仅改 import 行 + 锚点) |
| `runtime.py` | `inbound_verify/runtime.py` | 无(保留名) |
| `state_store.py` | `inbound_verify/state_store.py` | 无(**保留名**,仅改 import 行) |
| `expected_undelivered.py` | `inbound_verify/expected_undelivered.py` | 无(Tier1 保留名;Tier2 改名 compare) |
| `db_store.py` | `inbound_verify/store.py` | 无(叶子,无人 import) |
| `main_router.py` | `inbound_verify/cli/router.py` | 无(叶子) |
| `server.py` | `inbound_verify/cli/server.py` | 无(叶子) |
| `site_{shunxin,baishi,zto,yunda,anneng}.py` | `inbound_verify/sites/{…}.py`(去 `site_` 前缀) | 改 `runtime` + `main_router` 调用点;**grep 查残留引用兜底** |
> **命名策略(降低无测试下的风险)**:Tier 1 只对**被多处裸名引用**的模块(`state_store`、`expected_undelivered`、`runtime`、`paths`)**保留原名**,做到"仅改 import 行、零调用点改动";只对**叶子入口**(`db_store`/`main_router`/`server`)和**站点文件**(去 `site_` 前缀)做改名。`expected_undelivered` 的改名(`→ compare.py`)推迟到 Tier 2——那时本就要为抽 `domain.py` 重做该文件及其调用方,把改名 churn 并入一个已经在改的批次。这样 Tier 1 的纯搬运可被 `compileall` + 导入冒烟 + grep 充分验证,不依赖手测。
---
## 4. `paths.py` 锚点修正(唯一地雷,必须改对)
包搬进 `inbound_verify/` 后,`__file__` 多降一级。要把 `BASE_DIR` 继续指回项目根:
```python
# inbound_verify/paths.py
import os
# __file__ = .../InboundVerify/inbound_verify/paths.py
# 上两级 = .../InboundVerify (= 项目根,config.yaml/downloads/schema.sql 所在)
BASE_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
DOWNLOAD_DIR = os.path.join(BASE_DIR, "downloads")
OUTPUT_DIR = os.path.join(BASE_DIR, "output")
CONFIG_PATH = os.path.join(BASE_DIR, "config.yaml")
STATE_DB_PATH = os.path.join(BASE_DIR, "state", "state.db")
```
editable 安装不会移动文件,`__file__` 仍指向源码树,两级上跳稳定指向项目根。`store.py``SCHEMA_PATH = join(BASE_DIR, "schema.sql")` 自动跟着对。
---
## 5. import 改写规则
```python
from paths import ... from inbound_verify.paths import ...
import state_store from inbound_verify import state_store # 保留名,调用点 state_store.X 不变
from runtime import (...) from inbound_verify.runtime import (...)
import site_shunxin from inbound_verify.sites import shunxin # 5 站同理,调用点改 shunxin.X
import expected_undelivered from inbound_verify import expected_undelivered # Tier1 保留名
import expected_undelivered as eu from inbound_verify import expected_undelivered as eu # eu.X 不变
```
(Tier 2 后,`expected_undelivered` 改名 `compare`,`db_store` 的共享配置改 `from inbound_verify import domain`。)
---
## 6. 入口与打包
### 6.1 `main()` 包装(根目录不留薄壳)
```python
# inbound_verify/cli/server.py 末尾
def main():
uvicorn.run("inbound_verify.cli.server:app", host="0.0.0.0", port=8000)
if __name__ == "__main__":
main()
```
> `uvicorn.run` 改传**字符串** `"inbound_verify.cli.server:app"`(规范写法;不开 `reload`/`workers` 时仍在当前进程 import,行为等价)。`router.py` 套 `main()` 调 `run_multi_site_daemon()`;`store.py` 已有 `_cli`,改名为 `main`。
### 6.2 `pyproject.toml`(骨架)
```toml
[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"
[project]
name = "inbound-verify"
version = "0.1.0"
requires-python = ">=3.10"
dependencies = ["pandas>=2.0.0", "playwright>=1.40.0", "openpyxl>=3.1.0",
"PyYAML>=6.0", "websocket-client>=1.0.0", "fastapi>=0.110.0",
"uvicorn>=0.27.0", "apscheduler>=3.10.0", "psycopg[binary]>=3.1"]
[project.scripts]
inbound-verify = "inbound_verify.cli.router:main"
inbound-verify-server = "inbound_verify.cli.server:main"
inbound-verify-db = "inbound_verify.store:main"
[tool.setuptools.packages.find]
include = ["inbound_verify*"]
```
### 6.3 装包后三种启动方式(都汇到同一个 `main()`)
```bash
pip install -e . # 一次性
inbound-verify-server # 命令
python -m inbound_verify.cli.server # 兜底
uvicorn inbound_verify.cli.server:app # 生产最标准(端口/worker 命令行控)
```
---
## 7. 执行顺序与验证闸门(为"无测试"量身)
**安全原则**:Tier 1 必须先独立完成并验证为**行为等价**,再做任何动逻辑的改动。每段后必验证。验证职责分清——自动化部分我跑,端到端部分需你跑(真实登录/凭据/安能 Electron 只有你能提供)。
```
Tier 1 纯搬运(行为零改变)
├─ 我的自动验证: ① compileall 全过 ② 导入冒烟 ③ grep 查残留 site_ 引用 ④ DB 连通
└─ [等用户说"提交"] commit ← 安全基线
Tier 2 domain 抽取 / config 集中 / 路径消重 / expected_undelivered 改名 compare
├─ 我的自动验证: 同上
└─ [等用户说"提交"] commit
═══════ 手动测试闸门(你跑)═══════
全链路端到端:
起 inbound-verify → 登录 5 站(顺心双账号)→ 各站下载 → 比对(菜单 9)
→ inbound-verify-db ingest → 核对 output/应到未到数据.xlsx 与 PostgreSQL 三张表
通过? 否 → 修到通过
是 ↓
Tier 3 此时再定:anneng 拆不拆 / with_retry 抽不抽 base.py
└─ 每项后重跑我的自动验证,你按需复测
```
### 7.1 自动验证命令清单(我每个 Tier 后都跑)
```bash
.venv/Scripts/python.exe -m compileall inbound_verify # ① 语法
.venv/Scripts/python.exe -c "import inbound_verify.cli.router, \
inbound_verify.cli.server, inbound_verify.store, inbound_verify.runtime, \
inbound_verify.state_store" # ② 导入冒烟(三个入口会传递导入 sites/expected_undelivered 等;
# Tier2 改名后 expected_undelivered → compare,无需单独显式导入)
grep -rn "site_shunxin\|site_baishi\|site_zto\|site_yunda\|site_anneng" \
inbound_verify || echo "无残留 site_ 引用 ✓" # ③ 改名残留
.venv/Scripts/python.exe -c "from inbound_verify.store import _connect, _load_pg_config; \
c=_load_pg_config(); conn=_connect(c['dbname']); print('DB OK', conn.info.server_version)" # ④ DB
```
---
## 8. 各 Tier 内容
### Tier 1 — 纯搬运(行为零改变)
1. 建包骨架 + 空白 `__init__.py`(`sites/``cli/`)。
2. `git mv` §3.1 表中所有文件到新位置。
3. 按 §5 规则机械改写全部 import;站点改名后改 `runtime` + `main_router` 调用点(`shunxin.shunxin_expected_download(...)` 等)。
4.`paths.py` 锚点(§4)。
5. 三个入口加 `main()`(`store``_cli` 改名 `main`)。
6.`pyproject.toml`,`.venv``pip install -e .`
7. 跑 §7.1 四项自动验证。
8. 等用户说"提交"→ commit(子模块内)。
### Tier 2 — 定点小改进(每项后跑 §7.1)
- **抽 `domain.py`**:把 `expected_undelivered.py` 里的 `STATIONS`/`_site_cfg`/`ALL_REPORT_SITES`/`SITE_UNDELIVERED_FILE`/`BAISHI_FILE`/`BAISHI_COLUMNS`/`arrived_pieces_*` 移到 `inbound_verify/domain.py`;`store.py` 从依赖整个比对引擎改为 `from inbound_verify import domain`。**接缝最干净、收益明确,推荐做。**
- **`expected_undelivered.py``compare.py`** 改名,更新调用方(`store``as eu``runtime``cli/router`)。
- **加 `config.py`**:集中 `load_config()`(带缓存),各站点把自家的 `yaml.safe_load(open(CONFIG_PATH))` 换掉。
- **消重路径常量**:`compare.py` 自带那份 `BASE/DOWNLOADS/OUTPUT` 改成引用 `paths.py`
### Tier 3 — 大文件拆分(手动测试闸门之后;具体决策推迟到闸门)
- **`anneng.py``sites/anneng/` 子包**(`cdp.py`/`nav.py`/`expected.py`/`actual.py`):接缝分层清晰,但共享可变状态多(`CDP_PORT` 全局被 `set_cdp_port` 改、各种 URL hint、僵尸 tab 逻辑),CLAUDE.md 标注为脚gun。**倾向不拆**(1375 行虽大但是内聚的 CDP 驱动,强拆无测试网兜底风险高)——最终在闸门后定。
- **抽 `sites/base.py`**:`with_retry` / `_remove_if_exists` 在 5 站逐字重复;原作者在 CLAUDE.md 写"现阶段刻意不优化结构"。**抽不抽,闸门后定。**
---
## 9. 待决策(推迟到手动测试闸门)
| 决策 | 默认倾向 | 何时定 |
|---|---|---|
| `anneng.py` 拆不拆子包 | **不拆** | 手动测试通过后 |
| `with_retry`/`_remove_if_exists``base.py` | 待定 | 手动测试通过后 |
| `requirements.txt` 留还是删 | **留静态镜像**(注明 pyproject 为准) | 随时可改 |
---
## 10. 文档同步(Tier 1 必做)
- **README.md**:第二节目录树、第三节环境准备(加 `pip install -e .`)、第五节运行(改新命令)。
- **CLAUDE.md**:常用命令段全部改 `python -m inbound_verify…` / `inbound-verify…`,补 `pip install -e .``playwright install chromium``black`/`py_compile` 的包内路径写法。
---
## 11. 范围外
- **不做** src/ 布局。
- **不做** 给站点流程加 mock 测试。
- **不自动** commit/push(等用户明确指示)。
- 父仓库 `LogisticsHubIPA` 的子模块指针更新,是父仓库的单独一步,不在本 spec 范围。

View File

@@ -0,0 +1,290 @@
# 下载后自动入库钩子设计
- **日期**:2026-07-24
- **方案**:在 `runtime.dispatch_task` 下载成功分支挂一个**同步、kind 级、尽力而为**的入库钩子,调用新增的 `store.ingest_task(site, kind)`,仅入库本次刚下载的文件
- **力度**:后端最小闭环(钩子 + `ingest_task` + `ingest_state` 表 + `/status` 字段 + 配置开关/超时);不碰 dashboard UI、不上连接池/后台线程
- **状态**:已与用户对齐(4 个关键决策已逐项确认),待 spec 评审
---
## 1. 背景与目标
各站点的下载流程与 PostgreSQL 持久化模块(`store.py`)**都已实现**,但入库动作目前只能靠 CLI
(`inbound-verify-db ingest` / `python -m inbound_verify.store ingest`)手动触发,**没有挂到任何自动钩子上**。
本设计把入库动作挂到"数据完成下载之后"——具体挂在所有下载任务的唯一汇聚点
`runtime.dispatch_task` 的成功分支,与既有的"下载后写业务日期"钩子 `_record_business_date` 并列。
**目标**
1. 下载成功后**自动**把刚下载的那份数据 UPSERT 进 PostgreSQL无需手动跑 CLI。
2. 入库是**尽力而为**:失败只告警、绝不影响下载任务的成功判定(下载成功 = 任务成功)。
3. 入库结果可查:写入 `state_store`,经 `/api/status` 暴露,便于发现"入库坏了好几天"。
4. 可门控、可降级:配置开关 + PG 连接/语句超时,无 PG/cpolar 的开发机可整体跳过。
**非目标(本 spec 不做)**
- dashboard 前端展示 ingest 态(Next.js 侧API 字段已就绪待消费,单独小任务)。
- 4 站 `-未到数据.xlsx` 入库(既有设计仅百世未到入库4 站未到只给汇总报表)。
- 连接池、后台入库线程、失败补入重试(YAGNI)。
- 任何与下载流程本身、比对逻辑相关的改动。
---
## 2. 现状(决策依据)
### 2.1 汇聚点与既有钩子先例
- `runtime.dispatch_task(ctx, {"site","kind"})` 是全部 9 个下载任务 + 比对任务的**唯一执行入口**;
交互模式(`cli/router`)与服务模式(`cli/server` worker)都走它。
- 它的成功分支**已经有一个"下载后"钩子**:`_record_business_date(site, kind)`——写业务日期到
`state.db`best-effort,失败仅告警。新钩子天然挂在它旁边house style 完全一致。
- 4 站 `undelivered` 任务的 handler(`_site_undelivered_handler`)内部**直接调**
`TASK_HANDLERS[(site,"expected"/"actual")](ctx)`(不经 dispatch_task),故只触发**一次**外层
dispatch_task 钩子——`ingest_task(site,"undelivered")` 需据此同时入 expected+actual。
### 2.2 持久化模块现状
- `store.ingest(site=None)`:读 `downloads/` 现有 xlsx,幂等 UPSERT 到 PG。**站点级**:对顺心/中通/
韵达/安能同时入该站 expected+actual;百世只入 undelivered。**不感知 kind**。
- 三张表:`expected_record`(运单级) / `actual_record`(扫描件级) / `undelivered_record`(仅百世)。
- `_ingest_expected(cur, site, business_date)` / `_ingest_actual(cur, site)` /
`_ingest_undelivered_baishi(cur)` 为内部助手,直接复用,不改。
- `_connect(dbname)`:每次新建一条连接,`options=-c search_path=<schema>`。**经 cpolar 隧道**
(`5.tcp.cpolar.top:10364`),潜在延迟/抖动。
### 2.3 线程模型
- dispatch_task 跑在 Playwright 所属线程(router=主线程;server=单 worker 线程,串行消费
task_queue、空闲跑心跳)。**同步做 PG I/O 会阻塞这条线程**——故入库必须快、可超时、可降级。
### 2.4 数据流定位
- dashboard **不直接读 PG**:全部经 Next.js `/api/*` 代理到 InboundVerify FastAPI(读 state.db +
文件 + 汇总报表)。故自动入库的直接受益者是 **PG 这个下游数仓**(BI/长期归档/未来报表),
不是当前前端。这支撑了"同步内联、不上复杂调度"的判断。
### 2.5 state_store 现状(决定表设计)
- `site_status` 是**每站一行**、按 kind 展开列(`{kind}_ready/_generated_at/_business_date`)。
- `_upsert` 是**手写枚举列**的 read-modify-write(脆)。往里塞 ingest 列(4 列 × 3 kind = 12 列)
会很丑且易错——故选**独立 `ingest_state` 表**(用户选项里也提过"或一张很小的入库记录表")。
---
## 3. 四个关键决策(均已与用户确认)
| 决策 | 选定 | 理由 |
|---|---|---|
| 执行模型 | **同步内联** | 与 `_record_business_date` 一致、零新线程,贴合"刻意不优化结构"风格;cpolar 风险用超时+try/except 兜底 |
| 入库粒度 | **kind 级**(`ingest_task(site,kind)`) | 只入本次刚下载的文件,阻塞最小;"下了什么入什么";CLI `ingest(site)` 保留不动 |
| 配置门控 | **开关 + 连接超时** | `auto_ingest`(默认开)+ `connect_timeout_seconds`(默认 5);无 PG 开发机可关 |
| 失败可见性 | **stdout + state_store** | 仅 stdout 会在服务模式静默失败多天;落库后 `/api/status` 可查 |
---
## 4. 组件设计
### 4.1 `store.py` — 新增 `ingest_task(site, kind)` + 连接/配置增强
**新函数 `ingest_task(site, kind) -> int`**
单连接、单事务,复用现有 `_ingest_*`,返回总条数。路由表:
| site | kind | 调用 |
|---|---|---|
| `__compare__` | * | 无 → 返回 0 |
| 顺心/中通/韵达/安能 | `expected` | `_ingest_expected(cur, site, dates.get(site))` |
| 顺心/中通/韵达/安能 | `actual` | `_ingest_actual(cur, site)` |
| 顺心/中通/韵达/安能 | `undelivered` | `_ingest_expected` + `_ingest_actual` |
| 百世 | `undelivered` | `_ingest_undelivered_baishi(cur)` |
| 百世 | `expected`/`actual` | (无此任务)防御性返回 0 |
`dates = _read_business_dates()`(既有)。骨架:
```python
def ingest_task(site, kind):
"""按 (site, kind) 入库本次刚下载的文件(幂等 UPSERT)。返回总条数。
与 ingest(site) 的区别:只入本次刷新的那一类,避免重读写另一类文件。"""
if site == "__compare__":
return 0
dates = _read_business_dates()
total = 0
with _connect(_load_pg_config()["dbname"]) as conn:
with conn.cursor() as cur:
if kind == "expected":
total += _ingest_expected(cur, site, dates.get(site))
elif kind == "actual":
total += _ingest_actual(cur, site)
elif kind == "undelivered":
if site == "百世":
total += _ingest_undelivered_baishi(cur)
else: # 4 站:handler 内部连带下了 expected+actual
total += _ingest_expected(cur, site, dates.get(site))
total += _ingest_actual(cur, site)
# 其它组合(如 百世/expected):防御性 0
conn.commit()
return total
```
**`_load_pg_config()` 增键**:`auto_ingest`(默认 `True`)、`connect_timeout_seconds`(默认 `5`)。
**`_connect(dbname)` 增强**:
- `psycopg.connect(..., connect_timeout=c["connect_timeout_seconds"])`
- 连上后 `cur.execute("SET statement_timeout = '30s'")`(固定值,带注释说明可按需 knob 化;
connect_timeout 守隧道死连,statement_timeout 守失控查询,小 UPSERT 极少触发)。
- `create_database` / `init_schema` / `ingest` 共用此连接,30s 对它们无影响。
**新薄封装 `ingest_enabled() -> bool`**:读 `_load_pg_config()["auto_ingest"]`
供 runtime 钩子判定开关,避免钩子伸手进 store 私有函数。
**CLI `ingest-one`**(便于脱离下载单测路由):`python -m inbound_verify.store ingest-one 韵达 expected`
→ 直接调 `ingest_task(site, kind)``main()` 增该分支。
**不改**:既有 `ingest(site)` / `_ingest_*` / SQL / `domain`
### 4.2 `runtime.py` — 新增 `_persist_to_db(site, kind)`,挂到 `dispatch_task`
**`dispatch_task` 成功分支**(`_record_business_date` 之后)新增一行:
```python
ret = handler(ctx)
if ret is False:
return (state_store.TASK_FAILED, "任务执行失败(重试耗尽)")
_record_business_date(site, kind)
_persist_to_db(site, kind) # 新增:尽力而为,绝不外抛,不影响任务判定
return (state_store.TASK_SUCCESS, None)
```
**新函数 `_persist_to_db(site, kind) -> None`**:
- `store` **懒导入**(函数内 `from inbound_verify import store`),彻底回避 import 顺序/成环,冷路径无性能影响。
- `if site == "__compare__": return`
- 读开关:`if not store.ingest_enabled(): print(">> [入库] 已关闭,跳过"); return`
- 主逻辑(所有写库/写状态都包 try/except,绝不外抛):
```python
try:
count = store.ingest_task(site, kind)
state_store.set_ingest_state(site, kind, ok=True, count=count)
print(f">> [入库] {site}/{kind} 成功,{count}")
except Exception as e:
print(f">> [warn] 入库失败 {site}/{kind}: {e}")
try:
state_store.set_ingest_state(site, kind, ok=False, error=str(e))
except Exception as e2:
print(f">> [warn] 写入库状态也失败: {e2}")
```
### 4.3 `state_store.py` — 新增独立表 `ingest_state` + 读写
**`init_db()` 增表(幂等)**:
```sql
CREATE TABLE IF NOT EXISTS ingest_state (
site TEXT,
kind TEXT,
ok INTEGER, -- 1=成功 0=失败
ingested_at TEXT,
count INTEGER, -- 入库条数
error TEXT, -- 失败原因(成功则 '')
PRIMARY KEY (site, kind)
)
```
**新函数**:
- `set_ingest_state(site, kind, ok, count=0, error=None)`:UPSERT(短连接,`ingested_at=_now()`,
与现有 `set_*` 风格一致)。
- `get_all_ingest_state() -> {site: {kind: {ok, ingested_at, count, error}}}`(库不存在返回 `{}`)。
**心跳不刷 `ingest_state`**(只由钩子写)。澄清:复用的是 state_store 管线 + `/api/status` 通道,
不是心跳循环。**不碰 `site_status` / `_upsert`**。
### 4.4 `cli/server.py` — `/status` 暴露 ingest 态
`GET /status` 现有返回 `get_all_status()`;追加一个 `ingest` 键 = `state_store.get_all_ingest_state()`
一行改动,前端可按需消费。
### 4.5 配置 — `config.example.yaml` + `config.yaml`
`postgres` 段新增(两文件都加;`config.yaml` 已 gitignore):
```yaml
postgres:
# ...既有 host/port/user/password/dbname/schema...
auto_ingest: true # 下载成功后自动入库;false=跳过(无 PG/cpolar 的开发机)
connect_timeout_seconds: 5 # PG 连接超时(秒);cpolar 抖动兜底
```
顺手把 `config.yaml` 注释里过期的 `db_store.py` 改为 `store.py`(小清理)。
---
## 5. 数据流
```
dispatch_task 成功
→ _record_business_date(site, kind) # 既有:SQLite 业务日期,best-effort
→ _persist_to_db(site, kind) # 新增
auto_ingest? ─ no ─→ print 跳过,return
└ yes ─→ 懒导 store
store.ingest_task(site, kind)
_connect(connect_timeout) → SET statement_timeout
路由 _ingest_* → commit → 返回 count
成功 → set_ingest_state(ok,count) + print
异常 → set_ingest_state(fail,error) + warn
→ return TASK_SUCCESS(始终)
```
---
## 6. 错误处理矩阵
| 情形 | 行为 | 任务判定 |
|---|---|---|
| `auto_ingest=false` | print 跳过,return | SUCCESS(不变) |
| 文件缺失(下载刚成功却无文件,罕见) | `_ingest_*``[跳过]` 返回 0;ingest_task 返回 0;记 ok/count=0 | SUCCESS |
| PG 不可达 / connect_timeout / statement_timeout | psycopg 异常 → 捕获 → fail + warn | SUCCESS |
| 表未初始化(UndefinedTable) | 捕获 → fail + warn(提示"请先 `inbound-verify-db init`") | SUCCESS |
| `set_ingest_state` 自身失败 | 再包 try/except,绝不外抛(呼应 `_record_business_date`) | SUCCESS |
**核心不变式:入库的任何失败都不改变下载任务的成功判定。**
---
## 7. 测试
无 pytest(house 约定)。验证手段:
1. **路由单测(新 CLI)**:`.venv/Scripts/python.exe -m inbound_verify.store ingest-one 韵达 expected`
→ 确认只入韵达应到,不动韵达实到;`ingest-one 百世 undelivered` → 只入百世未到;
`ingest-one 顺心 undelivered` → 入顺心 expected+actual。
2. **端到端**:dashboard/API 触发一次下载 → 看 stdout `[入库]` 行 → 查 PG 行数 →
`GET /api/status``ingest` 字段反映 ok/count/ingested_at。
3. **降级**:`auto_ingest=false` → 下载仍 SUCCESS、无 `[入库]` 行;临时封掉 cpolar 端口 →
下载仍 SUCCESS、`[warn] 入库失败``ingest_state` 记 fail、`/status` 可见。
4. **回归**:`py_compile inbound_verify` + `black inbound_verify` + `compileall` 导入冒烟;
确认未改 `ingest(site)` 行为(手动跑一次 `store ingest` 全站入库仍正常)。
---
## 8. 已确认约束
- InboundVerify 是 git **子模块**;改动在子模块内,父仓库仅跟踪指针。
- **不加测试套件**;验证靠编译/导入冒烟 + 手动端到端。
- 改完 Python **必须跑 Black**
- **不自动提交/推送**:本仓库约定改动后等用户明确说"提交"再 commit/push(覆盖全局 auto-push 默认,
亦覆盖 brainstorming 默认的"写完即提交")。
---
## 9. 实现顺序提示(供 writing-plans 展开)
1. `store.py`:`_load_pg_config` 加键 + `_connect` 加超时/语句超时 → `ingest_task` → CLI `ingest-one`
2. `state_store.py`:`ingest_state` 建表 + `set_ingest_state` + `get_all_ingest_state`
3. `runtime.py`:`_persist_to_db` + 挂到 `dispatch_task`
4. `cli/server.py`:`/status``ingest` 字段。
5. 配置:`config.example.yaml` + `config.yaml` 加键 + 修过期注释。
6. Black + 编译/导入冒烟 + 手动端到端验证。

View File

@@ -80,7 +80,7 @@ Protocol error (Browser.setDownloadBehavior): Browser context management is not
```
- `type == "page"` 的才是普通页面(另有 `background_page``service_worker``webview` 等,按需过滤)。
- **`id` / `webSocketDebuggerUrl` 每次启动都会变**,所以一定要动态发现,**绝不能硬编码**(早期 `site_anneng.py` 硬编码 WS URL重启就失效
- **`id` / `webSocketDebuggerUrl` 每次启动都会变**,所以一定要动态发现,**绝不能硬编码**(早期 `sites/anneng.py` 硬编码 WS URL重启就失效
-`title``url` 来挑选你真正要操作的那一页。
### 3.2 CDP 消息往返
@@ -143,7 +143,7 @@ class CDP:
"""绑定到单个页面目标的同步 CDP 客户端。"""
def __init__(self, ws_url):
self.ws = websocket.create_connection(ws_url)
self.ws = websocket.create_connection(ws_url, timeout=15) # 15s socket 超时:防 Electron 业务 tab 偶发不回包时 recv 无限阻塞
self._id = 0
def call(self, method, **params):
@@ -408,8 +408,8 @@ def wait_for(self, selector, timeout=10, interval=0.3):
- 所有 Python 脚本在项目 `.venv` 内运行(`python -m venv .venv`,激活后安装依赖)。
- 依赖:`websocket-client`(已装)、`black`(格式化,已装)、可选 `pychrome`
- 相关脚本:`_probe_anneng.py`(探查)、`site_anneng.py`(菜单导航)、`download_clean.py`(免对话框下载)。
- 运行:`.venv/Scripts/python.exe download_clean.py`
- 相关脚本:`inbound_verify/sites/anneng.py`(菜单导航 + CDP 驱动;早期探查脚本 `_probe_anneng.py`、免对话框下载脚本 `download_clean.py` 已并入此模块,不再单独存在)。
- 运行:`.venv/Scripts/python.exe -m inbound_verify.sites.anneng expected|actual`
---

View File

@@ -1,7 +1,7 @@
# 后端各站点「应到件数 / 已到件数」统计逻辑审查报告
> 审查范围:`InboundVerify` 后端
> 核心模块:`expected_undelivered.py`(比对统计)、`runtime.py`(调度/下载路由)、`site_*.py`(各站数据采集)、`server.py`(对外接口)
> 核心模块:`compare.py`(比对统计)、`domain.py`(站点/文件/列配置)、`runtime.py`(调度/下载路由)、`sites/*.py`(各站数据采集)、`cli/server.py`(对外接口)
> 审查重点:每个站点的「应到件数」和「已到件数」如何统计、数据来源、口径与潜在歧义。
---
@@ -12,60 +12,60 @@
| 层 | 模块 | 职责 | 产物 |
|---|---|---|---|
| **① 数据采集层** | `site_顺心/中通/韵达/安能.py``site_百世.py` | 用 Playwright安能用 CDP登录各承运商后台导出 Excel 到 `downloads/` | `downloads/<站>-应到货物数据.xlsx``downloads/<站>-实到货物数据.xlsx``downloads/<站>-未到数据.xlsx` |
| **② 比对统计层** | `expected_undelivered.py` | 读 `downloads/` 源文件,两表比对,算出应到/到/未到件数并出报表 | `output/应到未到数据.xlsx`(汇总+各站明细) |
| **① 数据采集层** | `sites/{顺心,中通,韵达,安能,百世}.py` | 用 Playwright安能用 CDP登录各承运商后台导出 Excel 到 `downloads/` | `downloads/<站>-应到货物数据.xlsx``downloads/<站>-实到货物数据.xlsx``downloads/<站>-未到数据.xlsx` |
| **② 比对统计层** | `compare.py`(站点/文件/列配置取自 `domain.py` | 读 `downloads/` 源文件,两表比对,算出应到/到/未到件数并出报表 | `output/应到未到数据.xlsx`(汇总+各站明细) |
调度入口:`runtime.py``TASK_HANDLERS`
- 4 站:`expected` 下载 + `actual` 下载 + `undelivered`(先下应到+实到,再调 `write_site_file` 算出未到)。
- 百世:只有 `undelivered`(直接导「当日未扫」明细,无应到/实到两份基表)。
- `__compare__`(菜单[9]):调 `expected_undelivered.main()`,用 `downloads/` 现有文件生成全站汇总报表。
- `__compare__`(菜单[9]):调 `compare.main()`,用 `downloads/` 现有文件生成全站汇总报表。
---
## 二、核心统计公式4 站统一口径)
`expected_undelivered.process(name)` 是 4 站统计的唯一实现:
`compare.process(name)` 是 4 站统计的唯一实现:
```
对每一条应到运单(按运单号去重,保留首条):
应到件数 N = 应到数据中的「录单件数」(输出列名「总件数」
应到序号集 = {1, 2, …, N}
已到序号集 = 从「实到货物数据.xlsx」解析该运单实际到货的子单序号
未到序号集 = {1…N} 已到序号集
未到件数 += len(未到序号集) # 每缺 1 个子单 → 1 行未到明细
应到件数 += N
应到(按运单号去重,保留首条):
应到件数 N = 应到数据中的「交接件数」 # 注意:用「交接件数」,不是「录单件数」
(录单件数只是该单号总录单量,并非真正到站量)
应到件数 += N # 站点级应到 = Σ N
站点级
应到件数 = Σ N所有应到运单
未到件数 = Σ len(未到序号集)
已到件数 = 应到件数 未到件数 ← 注意:是「减出来」的,不是直接数实到
实到(直接数,不再由「应到−未到」倒推)
按「运单号」把实到表的「单号/子单号/扫描单号」分组,组内去重计数
→ 每个运单的实到件数;站点级实到 = Σ 各运单实到件数
未到:
未到件数 = max(0, 应到件数 实到件数) # 站点级
未到率 = 未到件数 ÷ 应到件数
完全未到运单 = 未到序号集 == {1…N} 的运单
部分未到运单 = 0 < 未到 < N 的运单
短少运单 = 实到件数 < 应到件数 N 的运单
(未到明细 downloads/<站>-未到数据.xlsx 只列这些短少运单
每行:交接单号 | 运单号 | 总件数(=N) | 已到单号1 | 已到单号2 | …)
```
**关键事实**`到件数` `应到件数 到件数` 推算出来的(`exp_pieces - len(rows)`),并非独立去数实到表。它等价于 `Σ|已到序号集|`,前提是实到解析正确且已到序号 ⊆ 应到序号
**关键事实**`到件数`直接数实到表「单号」、按运单号分组去重得到的(**不再由「应到−未到」倒推**`未到件数 = 应到件数 到件数`。前提是实到单号能正确按运单号分组(分组规则见下表各站 `arrived_pieces_*`
---
## 三、各站点统计明细
| 站点 | 应到件数来源 | 到件数来源 | 实到序号解析规则(`arrived` | 源文件(`downloads/` |
| 站点 | 应到件数来源 | 到件数来源 | 单号→运单号 分组规则(`arrived_pieces_*` | 源文件(`downloads/` |
|---|---|---|---|---|
| **中通** | `中通-应到货物数据.xlsx` 的「录单件数」 | 应到−未到(推算 | `运单号`是复合串 = 运单号 + 总数(4) + 顺序(4)取右 8 位,前 4 为总数、后 4 为顺序号 | `中通-应到货物数据.xlsx` / `中通-实到货物数据.xlsx` |
| **顺心** | `顺心-应到货物数据.xlsx` 的「录单件数」 | 应到−未到(推算 | `运单号`(干净) + `子单号` = 运单号 + 顺序号(3位);后缀(3位)即顺序号 | `顺心-应到货物数据.xlsx` / `顺心-实到货物数据.xlsx` |
| **韵达** | `韵达-应到货物数据.xlsx` 的「录单件数」 | 应到−未到(推算 | `主单号`=运单号`子单号`=主单号 + 顺序号(4位);后缀(4位)即顺序号 | `韵达-应到货物数据.xlsx` / `韵达-实到货物数据.xlsx` |
| **安能** | `安能-应到货物数据.xlsx` 的「录单件数」 | 应到−未到(推算 | `所属单号`=运单号`扫描单号`=所属单号 + 总数(4位) + 顺序号(4位)`has_total=True`,取末 4 位为顺序号 | `安能-应到货物数据.xlsx` / `安能-实到货物数据.xlsx` |
| **中通** | `中通-应到货物数据.xlsx` 的「交接件数」 | 实到表单号去重(直接数 | 实到`运单号`是复合串 = 运单号 + 总数(4) + 顺序(4)`v[:-8]` 归并到应到运单号,每条复合串即 1 件 | `中通-应到货物数据.xlsx` / `中通-实到货物数据.xlsx` |
| **顺心** | `顺心-应到货物数据.xlsx` 的「交接件数」 | 实到表单号去重(直接数 | `运单号`分组,`子单号`=每件(一件一个子单号) | `顺心-应到货物数据.xlsx` / `顺心-实到货物数据.xlsx` |
| **韵达** | `韵达-应到货物数据.xlsx` 的「交接件数」 | 实到表单号去重(直接数 | `主单号`分组`子单号`=每件 | `韵达-应到货物数据.xlsx` / `韵达-实到货物数据.xlsx` |
| **安能** | `安能-应到货物数据.xlsx` 的「交接件数」 | 实到表单号去重(直接数 | `所属单号`分组`扫描单号`=每件 | `安能-应到货物数据.xlsx` / `安能-实到货物数据.xlsx` |
| **百世** | **无**(站点只给未到) | **无(显示「—」)** | 不适用(无实到基表) | 仅 `百世-应到未到货物数据.xlsx`= 当日未扫明细,本身就是未到结果) |
> 4 站合计/图表口径:`expected_undelivered.build_summary` 只累加 4 站(`应到−实到`口径),**百世不计入合计**(无应到基数)。百世在表中单列,未到件数 = 其明细行数。
> 4 站合计/图表口径:`compare.build_summary` 只累加 4 站(`应到−实到`口径),**百世不计入合计**(无应到基数)。百世在表中单列,未到件数 = 其明细行数。
---
## 四、百世的特殊口径(务必注意)
百世是唯一「无应到/实到基数」的站点:
- 它的数据来自后台「扫描综合查询 → 到/接件扫描 → 当日 → 未扫」,导出即「当日未扫」明细(`site_baishi.py`)。
- 它的数据来自后台「扫描综合查询 → 到/接件扫描 → 当日 → 未扫」,导出即「当日未扫」明细(`sites/baishi.py`)。
- 因此 `process_baishi()` 只能给 `未到件数 = 行数``应到件数 = None``已到件数 = None`、完全/部分未到 = None。
- 报表里百世的应到/已到列显示「—」,未到率无法计算。
- **含义**:百世统计的是「今天还没扫到的件」,不是「相对应到总量的缺件率」。与 4 站口径不可直接相加比较。
@@ -87,7 +87,7 @@
| 接口 | 返回内容 | 是否含应到/已到计数 |
|---|---|---|
| `GET /status` | 各站 `login_state` + `expected/actual/undelivered_ready` + `business_date` + `worker_ready` | **不含**计数,只给就绪标志 |
| `GET /status` | 各站 `login_state` + `expected/actual/undelivered_ready` + `business_date` + `worker_ready` + `ingest`(每站每类入库 ok/count/时间) | **不含**应到/已到件数计数(`ingest` 是入库条数,非核销件数) |
| `GET /report` | `FileResponse(output/应到未到数据.xlsx)` | 计数只在 xlsx 里 |
| `GET /data/{filename}` | 下载 `downloads/` 下某源文件 | 原始数据,非统计值 |
@@ -97,8 +97,8 @@
## 七、潜在歧义与风险点(审查结论)
1. **已到件数是减出来的,不是数出来的**:依赖实到子单号解析正确。一旦某站 `子单号` 格式与解析规则(`arrived_*`)不匹配,该序号进不了「已到序号集」→ 误判为未到 → 已到件数被低估、未到率虚高。解析规则硬编码,承运商改版号段即失准。
2. **应到件数依赖「录单件数」**:若应到数据某运单 `录单件数` 缺失/为 0/非数字,该运单被 `continue` 跳过,既不计入应到也不计入未到 → 静默漏统(应到总量被低估,但该运单出现在实到中也因不在应到循环而永不计入已到,方向一致)。
1. **实到依赖单号→运单号分组正确**:实到件数靠把实到表「单号」按运单号分组去重计数(`arrived_pieces_*`)。一旦某站单号格式与分组规则不匹配(如复合串切分错),该件归不到对应运单 → 实到被低估、未到率虚高。规则硬编码,承运商改版号段即失准。
2. **应到件数依赖「交接件数」**:若应到数据某运单 `交接件数` 缺失/为 0/非数字,该运单被跳过,既不计入应到也不计入未到 → 静默漏统(应到总量被低估该运单即便出现在实到中也因不在应到循环而无处抵扣)。
3. **应到按运单号去重keep first**:同一运单多条交接记录只取首条 `录单件数`。若重复行的件数不同,取首条,可能与实际不符。
4. **百世不可并入合计**4 站合计的「已到总件数」不含百世;跨站看「已到」时别把百世当成有应到基数的站。
5. **应到/实到业务日期可能错位**(见第五节):两表不同步下载或偏移不一致时,比对口径失真。
@@ -111,9 +111,10 @@
| 文件 | 角色 |
|---|---|
| `expected_undelivered.py` | 统计核心:`process()`4站比对`process_baishi()`(百世)、`build_summary()`(汇总报表)`STATIONS`(各站解析配置) |
| `compare.py` | 统计核心:`process()`4站比对`process_baishi()`(百世)、`build_summary()`(汇总报表) |
| `domain.py` | 站点/文件名/列映射配置:`STATIONS`(各站解析配置)、`arrived_pieces_*`(实到单号→运单号分组)、`_site_cfg` |
| `runtime.py` | `TASK_HANDLERS`(下载/比对路由)、`_site_undelivered_handler`4站先下应到+实到再算未到)、`_record_business_date`(业务日期/偏移写入) |
| `site_中通/顺心/韵达/安能.py` | 各站 `expected_download` / `actual_download`Playwright/CDP 导出源表) |
| `site_百世.py` | `baishi_download_undelivered_data`(直接导「当日未扫」) |
| `sites/{中通,顺心,韵达,安能}.py` | 各站 `expected_download` / `actual_download`Playwright/CDP 导出源表) |
| `sites/baishi.py` | `baishi_download_undelivered_data`(直接导「当日未扫」) |
| `state_store.py` | `site_status`ready/business_date`site_config`offset/schedule`get_offset` |
| `server.py` | `/status``/report``/data/{filename}` 接口 |
| `cli/server.py` | `/status``/report``/data/{filename}` 接口 |

View File

@@ -1,11 +1,11 @@
# 韵达 / 安能 应到未到计算逻辑梳理(基于 expected_undelivered.py 重构版代码)
# 韵达 / 安能 应到未到计算逻辑梳理(基于 inbound_verify/compare.py 重构版代码)
> 用途:供审核两站「应到件数 / 实到件数 / 未到件数」的实现逻辑与关联键。
> 代码基准:`InboundVerify/expected_undelivered.py`(重构版)。行号对应本次读取。
> 代码基准:`InboundVerify/inbound_verify/compare.py`(重构版)。
---
## 、两站共用的计算骨架process 函数,第 134216 行
## 、两站共用的计算骨架process 函数)
无论韵达还是安能,最终都走同一个 `process(name)`

View File

@@ -0,0 +1 @@
# inbound_verify package

View File

@@ -0,0 +1 @@
# inbound_verify.cli — 入口(交互菜单 / 服务)

View File

@@ -1,8 +1,8 @@
# main_router.py
# inbound_verify/cli/router.py
#
# 交互菜单模式入口(调试 / 人工操作)。核心 Playwright 管理、任务派发、心跳
# 已抽到 runtime.py 共享;本文件只保留交互菜单与自动化测试。
# 服务模式(常驻 + FastAPI 接收指令)见 server.py。
# 服务模式(常驻 + FastAPI 接收指令)见 cli/server.py。
import os
import queue
@@ -11,8 +11,8 @@ import time
import yaml
from paths import CONFIG_PATH
from runtime import (
from inbound_verify.paths import CONFIG_PATH
from inbound_verify.runtime import (
APP_SITES,
HEARTBEAT_INTERVAL,
dispatch_task,
@@ -20,22 +20,18 @@ from runtime import (
run_heartbeat,
)
import state_store
from inbound_verify import state_store
# 各站点模块(自动化测试 + 比对用;任务派发在 runtime
import site_shunxin
import site_baishi
import site_zto
import site_yunda
import site_anneng
import expected_undelivered
from inbound_verify.sites import shunxin, baishi, zto, yunda, anneng
from inbound_verify import compare
def run_undelivered_compare():
"""应到未到比对(全站点):调用 expected_undelivered,读 downloads/ 下的应到/实到
"""应到未到比对(全站点):调用 compare读 downloads/ 下的应到/实到
数据生成 output/应到未到数据.xlsx汇总报表 + 各站明细"""
print("\n▶ 开始执行【应到未到比对(全站点)】任务 ...")
expected_undelivered.main()
compare.main()
# ====================================================================
@@ -57,20 +53,20 @@ def run_automation_test(pages_map):
# 不被模块内部"失败→重置→重试"机制掩盖。
flow_table = {
"顺心": {
"expected": ("应到", site_shunxin.shunxin_expected_download_impl),
"actual": ("实到", site_shunxin.shunxin_actual_download_impl),
"expected": ("应到", shunxin.shunxin_expected_download_impl),
"actual": ("实到", shunxin.shunxin_actual_download_impl),
},
"中通": {
"expected": ("应到", site_zto.zto_expected_download_impl),
"actual": ("实到", site_zto.zto_actual_download_impl),
"expected": ("应到", zto.zto_expected_download_impl),
"actual": ("实到", zto.zto_actual_download_impl),
},
"韵达": {
"expected": ("应到", site_yunda.yunda_expected_download_impl),
"actual": ("实到", site_yunda.yunda_actual_download_impl),
"expected": ("应到", yunda.yunda_expected_download_impl),
"actual": ("实到", yunda.yunda_actual_download_impl),
},
"安能": {
"expected": ("应到", site_anneng.anneng_expected_download_impl),
"actual": ("实到", site_anneng.anneng_actual_download_impl),
"expected": ("应到", anneng.anneng_expected_download_impl),
"actual": ("实到", anneng.anneng_actual_download_impl),
},
}
@@ -83,7 +79,7 @@ def run_automation_test(pages_map):
continue
for sx_idx, sx_page in enumerate(pages_map["顺心"], start=1):
try:
tag = site_shunxin.shunxin_belonging(sx_page)
tag = shunxin.shunxin_belonging(sx_page)
except Exception:
tag = f"账号{sx_idx}" # 读不到归属地时用序号占位,不阻断测试
for flow_key in CROSS_TEST_SEQUENCE:
@@ -102,7 +98,7 @@ def run_automation_test(pages_map):
(
"百世",
"应到未到",
site_baishi.baishi_download_undelivered_data_impl,
baishi.baishi_download_undelivered_data_impl,
pages_map["百世"],
"",
)
@@ -343,5 +339,10 @@ def run_multi_site_daemon():
print("程序已退出。")
if __name__ == "__main__":
def main():
"""交互菜单模式入口。"""
run_multi_site_daemon()
if __name__ == "__main__":
main()

View File

@@ -0,0 +1,435 @@
# server.py
#
# 服务模式入口FastAPI主线程uvicorn/asyncio+ Playwright worker独立线程
# 客户端经 HTTP 触发任务、查状态、下载数据worker 串行执行任务并跑心跳。
#
# 线程模型(关键):
# - 主线程uvicorn + FastAPI。路由【绝不】访问 Playwright 对象,只经
# task_queue投递任务+ state_store查状态/任务)+ 文件系统(下载数据)。
# - worker 线程runtime.launch_and_preparesync_playwright 在此)+ 任务循环,
# 独占所有 page 操作;与主线程仅经 Queue + SQLite 通信。
# 违反"路由不碰 Playwright"会崩sync 对象跨线程访问)。
#
# 运行python -m inbound_verify.cli.server (或 inbound-verify-server默认监听 0.0.0.0:8000
import os
import queue
import threading
import time
from contextlib import asynccontextmanager
from datetime import datetime, timedelta
from typing import Dict, Optional
import uvicorn
from apscheduler.schedulers.background import BackgroundScheduler
from apscheduler.triggers.interval import IntervalTrigger
from fastapi import FastAPI, HTTPException
from fastapi.responses import FileResponse
from pydantic import BaseModel
from inbound_verify.paths import OUTPUT_DIR
from inbound_verify import state_store
from inbound_verify.runtime import (
HEARTBEAT_INTERVAL,
TASK_HANDLERS,
dispatch_task,
launch_and_prepare,
run_heartbeat,
)
from inbound_verify import db_compare
# 全部站点;百世固定下载当天,不可配置偏移
ALL_SITES = ["顺心", "百世", "中通", "韵达", "安能"]
CONFIGURABLE_SITES = {"顺心", "中通", "韵达", "安能"}
# 任务队列:元素 (task_id, task_spec)。主线程投递worker 消费。
task_queue: "queue.Queue" = queue.Queue()
# worker 运行状态主线程只读worker 写)
worker_state = {
"ctx": None,
"stop": False,
"thread": None,
"ready": False, # launch_and_prepare 完成(各站就绪,可接任务)
"error": None, # worker 启动失败原因
}
# 每日定时下载调度器进程内job 只往 task_queue 投任务,不碰 Playwright
scheduler = BackgroundScheduler(daemon=True)
def _worker_loop():
"""worker 线程:启动 Playwright + 等就绪 + 任务循环(执行任务 + 心跳)。"""
try:
ctx = launch_and_prepare(foreground=False)
worker_state["ctx"] = ctx
worker_state["ready"] = True
# 【P1-2 重启自愈】worker 就绪后清理上轮遗留的 pending/running 僵尸任务
cleaned = state_store.fail_stale_tasks()
if cleaned:
print(
f">> [worker] 自愈:清理 {cleaned} 条遗留任务pending/running → failed"
)
print(">> [worker] 各站就绪,开始接收任务 ...")
except Exception as e:
worker_state["error"] = str(e)
print(f"❌ [worker] 启动失败: {e}")
return
last_heartbeat = 0.0
while not worker_state["stop"]:
try:
task_id, task_spec = task_queue.get(timeout=1)
except queue.Empty:
# 空闲时跑心跳
if time.monotonic() - last_heartbeat >= HEARTBEAT_INTERVAL:
run_heartbeat(ctx)
last_heartbeat = time.monotonic()
continue
state_store.update_task(task_id, state_store.TASK_RUNNING)
print(f">> [worker] 执行任务 #{task_id}: {task_spec}")
status, error = dispatch_task(ctx, task_spec)
state_store.update_task(task_id, status, error)
print(f">> [worker] 任务 #{task_id} 完成: {status} {error or ''}")
try:
ctx.stop()
except Exception:
pass
worker_state["ready"] = False
print(">> [worker] 已退出。")
def _in_active_window(active_start, active_end):
"""当前本地时间是否落在激活时段内(避免半夜空跑)。
- start/end 任一为空 → 不限时段24h 活跃)
- start == end非空→ 视为全天活跃
- start < end → 半开区间 [start, end)
- start > end → 跨午夜(如 22:00-06:00now >= start 或 now < end
"""
if not active_start or not active_end:
return True
now = datetime.now().strftime("%H:%M")
if active_start == active_end:
return True
if active_start < active_end:
return active_start <= now < active_end
return now >= active_start or now < active_end
def _enqueue_fetch(site, kind):
"""周期 job 回调worker 就绪 + 在激活时段内 + 该(site,kind)无未完成任务时,投递一次抓取任务。
跑在 APScheduler 线程池线程——绝不碰 Playwright只经 SQLite + task_queue 通信。"""
if not worker_state["ready"]:
return # worker 未就绪:与 POST /tasks 的 409 同义,下个周期补抓
cfg = state_store.get_fetch_schedule(site, kind)
if not cfg or not cfg["enabled"]:
return # 已禁用job 本应已注销,双重保险)
if not _in_active_window(cfg["active_start"], cfg["active_end"]):
return # 不在激活时段,跳过本次 fire
try:
target_date = state_store.resolve_target_date(site, kind)
tid = state_store.create_task_if_idle(
site, kind, trigger="auto", target_date=target_date
)
if tid is None:
return # 上一次同类任务还没跑完,跳过避免堆积
task_queue.put((tid, {"site": site, "kind": kind}))
print(f">> [周期] 投递 {site}/{kind} 任务 #{tid}")
except Exception as e:
print(f">> [周期] 投递 {site}/{kind} 失败: {e}")
def _reschedule_fetch(site, kind):
"""按持久化配置(重新)注册或取消该 (site,kind) 的周期抓取 jobIntervalTrigger"""
job_id = f"{site}_{kind}"
try:
scheduler.remove_job(job_id)
except Exception:
pass
cfg = state_store.get_fetch_schedule(site, kind)
if not cfg or not cfg["enabled"] or cfg["interval_minutes"] < 1:
return # 未启用或频率非法:不注册(等同取消)
try:
scheduler.add_job(
_enqueue_fetch,
IntervalTrigger(minutes=cfg["interval_minutes"]),
args=[site, kind],
id=job_id,
replace_existing=True,
)
print(
f">> [周期] 已注册 {site}/{kind}:每 {cfg['interval_minutes']} 分钟"
f"(激活 {cfg['active_start'] or '不限'}~{cfg['active_end'] or '不限'}"
)
except Exception as e:
print(f">> [周期] 注册 {site}/{kind} 失败: {e}")
@asynccontextmanager
async def lifespan(_app):
"""服务启停:起 worker 线程 / 通知 worker 停。"""
state_store.init_db() # 先建表/迁移状态库,确保早于 worker 就绪的 /api/status 可用
for site in ALL_SITES: # 按持久化配置注册各站各 kind 的周期抓取 job
for kind in state_store.allowed_kinds(site):
_reschedule_fetch(site, kind)
scheduler.start()
print(">> [定时] 调度器已启动")
t = threading.Thread(target=_worker_loop, daemon=True)
worker_state["thread"] = t
t.start()
yield
scheduler.shutdown(wait=False)
print(">> [定时] 调度器已停止")
worker_state["stop"] = True
t.join(timeout=10)
app = FastAPI(title="InboundVerify 服务端", lifespan=lifespan)
class TaskRequest(BaseModel):
site: str
kind: str
force: bool = (
False # 强制重下:忽略已落库去重,重新提交所有班次/交接单的导出任务(默认关)
)
date: Optional[str] = None # YYYY-MM-DD指定则下载该日数据否则走站点 offset
@app.post("/tasks")
def create_task(req: TaskRequest):
"""提交任务 {site, kind, force, date?} → 入队,返回 task_id。"""
# 【P0】后端未就绪时直接拒绝避免任务在 worker 启动前入队卡死
if not worker_state["ready"]:
raise HTTPException(
status_code=409, detail="后端尚未就绪,请等待各站点登录完成后再操作"
)
if (req.site, req.kind) not in TASK_HANDLERS:
raise HTTPException(status_code=400, detail=f"无效任务: {req.site}/{req.kind}")
# 指定日期合法性校验(仅在传了 date 时)
if req.date:
try:
target_date = datetime.strptime(req.date, "%Y-%m-%d").date()
except ValueError:
raise HTTPException(
status_code=400, detail=f"date 格式非法,需 YYYY-MM-DD: {req.date}"
)
today = datetime.now().date()
if target_date > today:
raise HTTPException(
status_code=400, detail=f"date 不可为未来日期: {req.date}"
)
if target_date < today - timedelta(days=31):
raise HTTPException(
status_code=400, detail=f"date 超出 31 天回溯上限: {req.date}"
)
if req.site == "百世":
raise HTTPException(
status_code=400, detail="百世固定下载当天,不支持指定日期"
)
task_id = state_store.create_task(
req.site,
req.kind,
trigger="manual",
target_date=state_store.resolve_target_date(req.site, req.kind, req.date),
force=req.force,
)
spec = {"site": req.site, "kind": req.kind, "force": req.force}
if req.date:
spec["date"] = req.date
task_queue.put((task_id, spec))
return {"task_id": task_id}
@app.get("/tasks/{task_id}")
def get_task(task_id: int):
t = state_store.get_task(task_id)
if not t:
raise HTTPException(status_code=404, detail="任务不存在")
return t
@app.get("/tasks")
def list_tasks(limit: int = 20):
return state_store.list_tasks(limit)
# ── DB 比对(基于 PostgreSQL不依赖 Excel 文件)──
class CompareRequest(BaseModel):
site: str
date: str # YYYY-MM-DD
@app.post("/compare")
def run_compare(req: CompareRequest):
"""DB 差缺比对:以实到扫描日期为锚点,反推交接批次,展开全量比对。
返回统计指标 + 差缺明细。
"""
# 合法性校验
if req.site not in db_compare.SITE_COMPARE_CONFIG:
raise HTTPException(
status_code=400,
detail=f"不支持的站点: {req.site}(支持: {list(db_compare.SITE_COMPARE_CONFIG.keys())}",
)
try:
target_date = datetime.strptime(req.date, "%Y-%m-%d").date()
except ValueError:
raise HTTPException(
status_code=400, detail=f"date 格式非法,需 YYYY-MM-DD: {req.date}"
)
today = datetime.now().date()
if target_date > today:
raise HTTPException(status_code=400, detail=f"date 不可为未来日期: {req.date}")
result = db_compare.compare_site_date(req.site, req.date)
if result is None:
raise HTTPException(
status_code=404,
detail=f"{req.site} {req.date}: 当天无实到数据,无法比对",
)
return {
"site": result.site,
"date": result.date,
"batches": result.batches,
"stats": {
"waybill_count": result.stats.waybill_count,
"sf_wb_count": result.stats.sf_wb_count,
"expected_pieces": result.stats.expected_pieces,
"arrived_pieces": result.stats.arrived_pieces,
"undelivered_pieces": result.stats.undelivered_pieces,
"undelivered_wb": result.stats.undelivered_wb,
"full_miss": result.stats.full_miss,
"part_miss": result.stats.part_miss,
"sf_undelivered": result.stats.sf_undelivered,
},
"rows": [
{
"handover_no": r.handover_no,
"waybill_no": r.waybill_no,
"total_pieces": r.total_pieces,
"arrived_pieces": r.arrived_pieces,
"arrived_list": r.arrived_list,
"is_sf": r.is_sf,
}
for r in result.rows
],
}
@app.get("/status")
def get_status():
"""各站登录态 + 数据态 + 入库态(前端状态盘用),另含 worker 就绪状态。"""
return {
"worker_ready": worker_state["ready"],
"worker_error": worker_state["error"],
"sites": state_store.get_all_status(),
"ingest": state_store.get_all_ingest_state(),
}
class FetchScheduleSpec(BaseModel):
enabled: bool
active_start: str = ""
active_end: str = ""
interval_minutes: int
class ConfigRequest(BaseModel):
expected_offset: Optional[int] = None
actual_offset: Optional[int] = None
# DEPRECATED旧"每日定点"字段,保留一个发布周期兼容旧前端请求(收到即忽略)
schedule_enabled: Optional[bool] = None
schedule_time: Optional[str] = None
settings: Optional[Dict[str, str]] = None
fetch_schedules: Optional[Dict[str, FetchScheduleSpec]] = None
def _default_cfg():
return {
"expected_offset": 0,
"actual_offset": 0,
"fetch_schedules": {},
}
@app.get("/config")
def get_config():
"""各站配置:应到/实到偏移 + 每日定时(百世偏移恒 0"""
cfg = state_store.get_all_config()
return {site: cfg.get(site, _default_cfg()) for site in ALL_SITES}
@app.put("/config/{site}")
def set_config(site: str, req: ConfigRequest):
if site not in ALL_SITES:
raise HTTPException(status_code=400, detail=f"未知站点: {site}")
# 应到/实到偏移(百世锁定当天)
for kind, val in (
("expected", req.expected_offset),
("actual", req.actual_offset),
):
if val is not None:
if site == "百世":
raise HTTPException(
status_code=400, detail="百世固定下载当天,不可配置偏移"
)
state_store.set_offset(site, kind, val)
# 周期抓取调度(应到/实到各自独立;百世只允许 undelivered
if req.fetch_schedules:
allowed = state_store.allowed_kinds(site)
for kind, spec in req.fetch_schedules.items():
if kind not in allowed:
raise HTTPException(
status_code=400,
detail=f"站点 {site} 不支持抓取类型 {kind}(允许: {list(allowed)}",
)
interval = max(1, min(1440, int(spec.interval_minutes)))
state_store.set_fetch_schedule(
site,
kind,
spec.enabled,
spec.active_start,
spec.active_end,
interval,
)
_reschedule_fetch(site, kind)
# 站点专属配置(密码/账号/路径…)
if req.settings:
for k, v in req.settings.items():
state_store.set_setting(site, k, v)
cfg = state_store.get_all_config().get(site, _default_cfg())
cfg["settings"] = state_store.get_site_settings(site)
return cfg
@app.get("/config/{site}/settings")
def get_site_settings_api(site: str):
"""单站专属配置(密码/账号/路径…;不进 5s 轮询)。"""
if site not in ALL_SITES:
raise HTTPException(status_code=400, detail=f"未知站点: {site}")
return state_store.get_site_settings(site)
REPORT_FILE = "应到未到数据.xlsx"
@app.get("/report")
def download_report():
"""下载 output/应到未到数据.xlsx跑比对生成未生成 404"""
path = os.path.join(OUTPUT_DIR, REPORT_FILE)
if not os.path.isfile(path):
raise HTTPException(status_code=404, detail="报告尚未生成,请先跑比对")
return FileResponse(path, filename=REPORT_FILE)
def main():
"""服务模式入口。传字符串导入路径(规范写法;不开 reload/workers 时进程内 import行为等价"""
uvicorn.run("inbound_verify.cli.server:app", host="0.0.0.0", port=8000)
if __name__ == "__main__":
main()

View File

@@ -40,95 +40,22 @@ from openpyxl.chart import BarChart, Reference
from openpyxl.worksheet.page import PageMargins
from openpyxl.worksheet.properties import PageSetupProperties
BASE = os.path.dirname(os.path.abspath(__file__))
DOWNLOADS = os.path.join(BASE, "downloads")
OUTPUT = os.path.join(BASE, "output")
OUTFILE = os.path.join(OUTPUT, "应到未到数据.xlsx")
from inbound_verify.paths import DOWNLOAD_DIR, OUTPUT_DIR
from inbound_verify.domain import (
ALL_REPORT_SITES,
BAISHI_COLUMNS,
BAISHI_FILE,
SITE_UNDELIVERED_FILE,
STATIONS,
_site_cfg,
arrived_pieces_by_cols,
arrived_pieces_zhongtong,
)
# 汇总报表覆盖的全部站点4 站在前、百世在末;汇总页图表只取 4 站)
ALL_REPORT_SITES = ["顺心", "中通", "韵达", "安能", "百世"]
# 4 站单站未到明细文件名(百世未到文件由站点直接产出,名为 BAISHI_FILE
SITE_UNDELIVERED_FILE = "{name}-未到数据.xlsx"
BAISHI_FILE = "百世-应到未到货物数据.xlsx"
BAISHI_COLUMNS = ["类型", "子单号", "运单号", "最新扫描记录"]
# 比对报表输出文件(路径锚定统一走 paths.py
OUTFILE = os.path.join(OUTPUT_DIR, "应到未到数据.xlsx")
# ============================ 比对逻辑 ============================
def arrived_pieces_zhongtong(df):
"""中通实到「运单号」为复合串H + 运单号(12) + 总数(4) + 顺序(4))。
基号 = v[:-8]与应到表运单号对齐单件 = 整串每串即一件"""
res = defaultdict(set)
for v in df["运单号"]:
v = str(v).strip()
if len(v) > 8 and v[-4:].isdigit():
res[v[:-8]].add(v) # 以完整复合串作为“已到单号”存入
return res
def arrived_pieces_by_cols(wb_col, piece_col):
"""顺心 / 韵达 / 安能:按干净运单列分组,单件 = 子单号 / 扫描单号。
wb_col实到表中与应到运单号对齐的干净列
顺心=运单号 / 韵达=主单号 / 安能=所属单号
piece_col实到表中每件货物的单号列子单号 / 扫描单号"""
def parse(df):
res = defaultdict(set)
for m, s in zip(df[wb_col], df[piece_col]):
m, s = str(m).strip(), str(s).strip()
if m and s:
res[m].add(s)
return res
return parse
STATIONS = [
{
"name": "中通",
"exp": "中通-应到货物数据.xlsx",
"act": "中通-实到货物数据.xlsx",
"exp_qty": "交接件数", # 应到件数口径:交接件数(非录单件数)
"exp_wb": "运单号", # 应到表运单号列(兼作去重键)
"exp_jd": "交接单号", # 未到数据需展示的交接单号
"arrived_pieces": arrived_pieces_zhongtong,
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "顺心",
"exp": "顺心-应到货物数据.xlsx",
"act": "顺心-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("运单号", "子单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "韵达",
"exp": "韵达-应到货物数据.xlsx",
"act": "韵达-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("主单号", "子单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "安能",
"exp": "安能-应到货物数据.xlsx",
"act": "安能-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("所属单号", "扫描单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
]
def _site_cfg(name):
"""按名称取 4 站配置(百世不在 STATIONS返回 None"""
return next((c for c in STATIONS if c["name"] == name), None)
# 站点 / 文件名 / 列映射配置ALL_REPORT_SITES / STATIONS / _site_cfg / BAISHI_FILE 等)见 domain.py。
def process(name):
@@ -139,8 +66,8 @@ def process(name):
cfg = _site_cfg(name)
if cfg is None:
return None
exp_path = os.path.join(DOWNLOADS, cfg["exp"])
act_path = os.path.join(DOWNLOADS, cfg["act"])
exp_path = os.path.join(DOWNLOAD_DIR, cfg["exp"])
act_path = os.path.join(DOWNLOAD_DIR, cfg["act"])
if not os.path.exists(exp_path) or not os.path.exists(act_path):
print(f"[跳过] {cfg['name']}downloads 下缺少 {cfg['exp']}{cfg['act']}")
return None
@@ -148,6 +75,12 @@ def process(name):
df_exp = pd.read_excel(exp_path, dtype=str).fillna("")
df_act = pd.read_excel(act_path, dtype=str).fillna("")
if name == "韵达":
# 韵达实到数据有重复行(同子单号出现两次),保留交接单号为空的(到/接件扫描),
# 丢弃交接单号不为空的(派件/签收等),再按子单号去重。
df_act = df_act[df_act["交接单号"].astype(str).str.strip() == ""]
df_act = df_act.drop_duplicates(subset=["子单号"], keep="last")
# 同一运单可能有多条交接记录,按运单号去重、保留首条
dup = int(df_exp[cfg["exp_wb"]].duplicated().sum())
df_exp = df_exp.drop_duplicates(subset=[cfg["exp_wb"]], keep="first")
@@ -291,7 +224,7 @@ def write_station(ws, columns, rows):
def process_baishi():
"""百世:读站点直供的未到明细,返回 (columns, rows, stats);文件缺失返回 None。
百世文件本身即未到结果无应到/已到基数统计只能给出未到件数"""
path = os.path.join(DOWNLOADS, BAISHI_FILE)
path = os.path.join(DOWNLOAD_DIR, BAISHI_FILE)
if not os.path.exists(path):
return None
df = pd.read_excel(path, dtype=str).fillna("")
@@ -301,7 +234,9 @@ def process_baishi():
# 应到/实到基数取自「扫描综合查询」应扫/已扫(到/接件扫描→当日),
# 由 baishi_download_undelivered_data_impl 在同次导航里抓取并落 site_settings。
# 未抓取过则 get_setting 返回 "" → 视为无基数(报表显示「—」)。
import state_store # 与 _read_business_dates 一致:比对模块纯离线,懒加载
from inbound_verify import (
state_store,
) # 与 _read_business_dates 一致:比对模块纯离线,懒加载
def _to_int(v):
v = (v or "").strip().replace(",", "")
@@ -329,7 +264,7 @@ def process_baishi():
def write_site_file(name):
"""4 站:把该站未到明细写到 downloads/<站>-未到数据.xlsx。
应到/实到缺process 返回 None 删旧文件返回 False成功返回 True"""
path = os.path.join(DOWNLOADS, SITE_UNDELIVERED_FILE.format(name=name))
path = os.path.join(DOWNLOAD_DIR, SITE_UNDELIVERED_FILE.format(name=name))
out = process(name)
if out is None:
if os.path.exists(path):
@@ -348,7 +283,7 @@ def _read_business_dates(include):
"""从状态库读各站业务日期dispatch 下载成功时快照写入),供报告「数据日期」列。
4 站取 expected_business_date报告按应到口径百世取 undelivered_business_date
从未下过的站返回空串诚实留空不反推"""
import state_store # lazy import比对模块本身保持纯离线
from inbound_verify import state_store # lazy import比对模块本身保持纯离线
status = state_store.get_all_status()
dates = {}
@@ -365,7 +300,7 @@ def build_full_report(include, dates=None):
"""生成全站汇总报表 output/应到未到数据.xlsx。
include: 本次成功的站点集合未成功站点在汇总里保留行无数据不影响他站
返回 {站点: 未到件或None} 供日志"""
os.makedirs(OUTPUT, exist_ok=True)
os.makedirs(OUTPUT_DIR, exist_ok=True)
wb = Workbook()
wb.remove(wb.active)
summary_ws = wb.create_sheet("汇总报表") # 首页占位
@@ -534,8 +469,16 @@ def build_summary(ws, results, generated_at, dates=None):
# 百世:未到明细已知;若已抓取应到/实到基数(扫描综合查询应扫/已扫)则填真实值
if s["应到件"] is not None and s["已到件"] is not None:
srate = (s["未到件"] / s["应到件"]) if s["应到件"] else 0
vals = [name, s["运单数"], s["应到件"], s["已到件"],
s["未到件"], srate, "", ""]
vals = [
name,
s["运单数"],
s["应到件"],
s["已到件"],
s["未到件"],
srate,
"",
"",
]
else:
vals = [name, s["运单数"], "", "", s["未到件"], "", "", ""]
else:
@@ -564,7 +507,12 @@ def build_summary(ws, results, generated_at, dates=None):
cell.fill = PatternFill("solid", fgColor=ZEBRA)
if isinstance(v, (int, float)):
cell.number_format = "0.0%" if i == 5 else "#,##0"
if i == 5 and s is not None and isinstance(v, (int, float)) and not isinstance(v, bool):
if (
i == 5
and s is not None
and isinstance(v, (int, float))
and not isinstance(v, bool)
):
cell.fill = PatternFill("solid", fgColor=heat(srate))
ws.row_dimensions[r].height = 19
r += 1
@@ -647,14 +595,14 @@ def main():
include = set()
for name in ALL_REPORT_SITES:
if name == "百世":
if os.path.exists(os.path.join(DOWNLOADS, BAISHI_FILE)):
if os.path.exists(os.path.join(DOWNLOAD_DIR, BAISHI_FILE)):
include.add(name)
else:
cfg = _site_cfg(name)
if (
cfg
and os.path.exists(os.path.join(DOWNLOADS, cfg["exp"]))
and os.path.exists(os.path.join(DOWNLOADS, cfg["act"]))
and os.path.exists(os.path.join(DOWNLOAD_DIR, cfg["exp"]))
and os.path.exists(os.path.join(DOWNLOAD_DIR, cfg["act"]))
):
include.add(name)
if not include:

View File

@@ -0,0 +1,658 @@
# -*- coding: utf-8 -*-
"""
db_compare.py — 基于 PostgreSQL 的应到未到差缺比对引擎。
与 compare.pyExcel 版)并行:本模块直接从 DB 查询数据进行比对,
不依赖 downloads/ 下的 Excel 文件。
核心思路:以实到扫描日期为锚点 → 反推交接批次 → 展开批次全量比对。
每个站点只需提供配置waybill 列名 / piece 列名 / 是否有 SF 特殊处理),
核心比对逻辑完全通用。
顺心站点 SF 运单特殊处理SF 运单的子单号piece_no为随机号码不能用
COUNT(DISTINCT piece_no) 去重计数,改为 COUNT(*) 行计数。
用法:
from inbound_verify.db_compare import compare_site_date, SITE_COMPARE_CONFIG
result = compare_site_date("顺心", "2026-07-25")
if result:
print(result.stats)
for row in result.rows:
print(row)
"""
import os
from dataclasses import dataclass, field
from datetime import date, datetime, timedelta
import psycopg
import yaml
from openpyxl import Workbook
from openpyxl.styles import Font, PatternFill, Alignment, Border, Side
from inbound_verify.paths import CONFIG_PATH, OUTPUT_DIR, DOWNLOAD_DIR
from inbound_verify.domain import _site_cfg, ALL_REPORT_SITES, BAISHI_COLUMNS
# ============================== 结果类型 ==============================
@dataclass
class CompareStats:
"""单站点/单批次比对统计。"""
waybill_count: int = 0 # 应到运单数
expected_pieces: int = 0 # 应到件数
arrived_pieces: int = 0 # 实到件数
undelivered_pieces: int = 0 # 未到件数
undelivered_wb: int = 0 # 差缺运单数
full_miss: int = 0 # 完全未到
part_miss: int = 0 # 部分未到
sf_wb_count: int = 0 # SF 运单数
sf_undelivered: int = 0 # SF 差缺数
@dataclass
class UndeliveredRow:
"""单条差缺明细。"""
handover_no: str = "" # 交接单号
waybill_no: str = "" # 运单号
total_pieces: int = 0 # 总件数(交接件数)
arrived_pieces: int = 0 # 已到件数
arrived_list: list = field(default_factory=list) # 已到单号列表
is_sf: bool = False # 是否 SF 运单
@dataclass
class CompareResult:
"""一次比对的完整结果。"""
site: str = ""
date: str = ""
batches: list = field(default_factory=list) # 涉及的交接批次
stats: CompareStats = field(default_factory=CompareStats)
rows: list = field(default_factory=list) # UndeliveredRow 列表
# ============================== 站点比对配置 ==============================
@dataclass
class SiteCompareConfig:
"""DB 比对的站点参数。"""
name: str # 站点名
has_sf: bool = False # 是否需要区分 SF 运单
# 四站点 DB 比对配置(百世不参与 4 站比对)
SITE_COMPARE_CONFIG: dict[str, SiteCompareConfig] = {
"顺心": SiteCompareConfig(name="顺心", has_sf=True),
"中通": SiteCompareConfig(name="中通", has_sf=False),
"韵达": SiteCompareConfig(name="韵达", has_sf=False),
"安能": SiteCompareConfig(name="安能", has_sf=False),
}
# ============================== DB 连接 ==============================
def _load_pg_config():
"""从 config.yaml 读 postgres 段。与 store.py 共用同一配置源。"""
if not os.path.exists(CONFIG_PATH):
raise FileNotFoundError(
f"未找到配置文件 {CONFIG_PATH}(请参考 config.example.yaml 创建 config.yaml"
)
with open(CONFIG_PATH, "r", encoding="utf-8") as f:
cfg = yaml.safe_load(f) or {}
pg = cfg.get("postgres") or {}
return {
"host": pg.get("host", "127.0.0.1"),
"port": int(pg.get("port", 5432)),
"user": pg.get("user", "postgres"),
"password": pg.get("password", ""),
"dbname": pg.get("dbname", "CQHXDB"),
"schema": pg.get("schema", "inbound_verify"),
"connect_timeout_seconds": int(pg.get("connect_timeout_seconds", 5)),
}
def _connect():
c = _load_pg_config()
return psycopg.connect(
host=c["host"],
port=c["port"],
dbname=c["dbname"],
user=c["user"],
password=c["password"],
options=f"-c search_path={c['schema']} -c statement_timeout=30s",
connect_timeout=c["connect_timeout_seconds"],
)
# ============================== 核心比对逻辑 ==============================
def compare_site_date(site: str, target_date: str) -> CompareResult | None:
"""对指定站点和日期执行 DB 差缺比对。
算法:
1. 取 scan_time::date = target_date 的实到运单(锚点)
2. 反推这些运单所属的交接批次handover_no
3. 展开批次全量应到运单
4. 查询批次全量实到扫描
5. 逐运单比对差缺SF/non-SF 分支处理)
Args:
site: 站点名("顺心"/"中通"/"韵达"/"安能"
target_date: 日期 "YYYY-MM-DD"
Returns:
CompareResult 或 None当天无实到数据时返回 None
"""
cfg = SITE_COMPARE_CONFIG.get(site)
if cfg is None:
print(f"[db_compare] 不支持的站点: {site}")
return None
try:
conn = _connect()
cur = conn.cursor()
# ── Step 1: 取实到锚点 ──
cur.execute(
"""
SELECT DISTINCT waybill_no FROM actual_record
WHERE site = %s AND scan_time::date = %s
""",
(site, target_date),
)
anchor_wbs = [r[0] for r in cur.fetchall()]
if not anchor_wbs:
print(f"[db_compare] {site} {target_date}: 当天无实到数据")
conn.close()
return None
# ── Step 2: 反推交接批次 ──
cur.execute(
"""
SELECT DISTINCT e.handover_no FROM expected_record e
WHERE e.site = %s AND e.waybill_no = ANY(%s)
""",
(site, anchor_wbs),
)
batches = [r[0] for r in cur.fetchall()]
# ── Step 3: 展开批次全量应到 ──
cur.execute(
"""
SELECT waybill_no, handover_no, handover_pieces
FROM expected_record
WHERE site = %s AND handover_no = ANY(%s)
ORDER BY handover_no, waybill_no
""",
(site, batches),
)
exp_rows = cur.fetchall() # [(waybill_no, handover_no, handover_pieces), ...]
if not exp_rows:
conn.close()
return None
all_wbs = [r[0] for r in exp_rows]
# ── Step 4: 取批次全量实到 ──
cur.execute(
"""
SELECT waybill_no, piece_no FROM actual_record
WHERE site = %s AND waybill_no = ANY(%s)
ORDER BY waybill_no, piece_no
""",
(site, all_wbs),
)
act_rows = cur.fetchall() # [(waybill_no, piece_no), ...]
conn.close()
# ── Step 5: 逐运单比对 ──
return _do_compare(site, target_date, batches, exp_rows, act_rows, cfg)
except Exception as e:
print(f"[db_compare] {site} {target_date} 比对异常: {e}")
return None
def compare_site_batch(site: str, handover_no: str) -> CompareResult | None:
"""按指定交接单号执行全批次比对(不依赖实到锚点)。
用于已知交接单号后精确比对某一批次。
"""
cfg = SITE_COMPARE_CONFIG.get(site)
if cfg is None:
print(f"[db_compare] 不支持的站点: {site}")
return None
try:
conn = _connect()
cur = conn.cursor()
cur.execute(
"""
SELECT waybill_no, handover_no, handover_pieces
FROM expected_record
WHERE site = %s AND handover_no = %s
ORDER BY waybill_no
""",
(site, handover_no),
)
exp_rows = cur.fetchall()
if not exp_rows:
conn.close()
return None
all_wbs = [r[0] for r in exp_rows]
cur.execute(
"""
SELECT waybill_no, piece_no FROM actual_record
WHERE site = %s AND waybill_no = ANY(%s)
ORDER BY waybill_no, piece_no
""",
(site, all_wbs),
)
act_rows = cur.fetchall()
conn.close()
return _do_compare(
site,
f"batch:{handover_no}",
[handover_no],
exp_rows,
act_rows,
cfg,
)
except Exception as e:
print(f"[db_compare] {site} batch:{handover_no} 比对异常: {e}")
return None
# ============================== 比对核心 ==============================
def _do_compare(
site: str,
label: str,
batches: list[str],
exp_rows: list[tuple], # [(waybill_no, handover_no, handover_pieces), ...]
act_rows: list[tuple], # [(waybill_no, piece_no), ...]
cfg: SiteCompareConfig,
) -> CompareResult:
"""执行逐运单比对,产出统计 + 差缺明细。
与 compare.py:process() 口径一致:
- 应到件数 = handover_pieces交接件数
- 实到件数 = SF ? COUNT(*) : COUNT(DISTINCT piece_no)
- arrived_cnt >= handover_pieces → 足额到货,跳过
"""
# 构建实到索引: waybill_no → [piece_no, ...](保留所有行,不去重)
act_by_wb: dict[str, list[str]] = {}
for wb, piece in act_rows:
act_by_wb.setdefault(wb, []).append(piece)
stats = CompareStats()
rows: list[UndeliveredRow] = []
max_arrived = 0
for wb, handover_no, handover_pcs in exp_rows:
handover_pcs = handover_pcs or 0
if handover_pcs <= 0:
continue
stats.waybill_count += 1
stats.expected_pieces += handover_pcs
is_sf = cfg.has_sf and wb.startswith("SF")
if is_sf:
stats.sf_wb_count += 1
all_pieces = act_by_wb.get(wb, [])
if is_sf:
# SF: 行计数不去重piece_no 是随机号码)
arrived_cnt = len(all_pieces)
arrived_list = list(all_pieces)
else:
# non-SF: 子单号去重
unique_pieces = list(dict.fromkeys(all_pieces)) # 保序去重
arrived_cnt = len(unique_pieces)
arrived_list = unique_pieces
stats.arrived_pieces += arrived_cnt
if arrived_cnt >= handover_pcs:
continue # 足额或溢到,不进差缺表
if arrived_cnt == 0:
stats.full_miss += 1
else:
stats.part_miss += 1
if is_sf:
stats.sf_undelivered += 1
max_arrived = max(max_arrived, arrived_cnt)
rows.append(
UndeliveredRow(
handover_no=handover_no,
waybill_no=wb,
total_pieces=handover_pcs,
arrived_pieces=arrived_cnt,
arrived_list=arrived_list,
is_sf=is_sf,
)
)
stats.undelivered_pieces = max(0, stats.expected_pieces - stats.arrived_pieces)
stats.undelivered_wb = stats.full_miss + stats.part_miss
result = CompareResult(
site=site,
date=label,
batches=batches,
stats=stats,
rows=rows,
)
# 打印摘要
print(
f"[db_compare] {site} {label}: "
f"batches={len(batches)}, "
f"wb={stats.waybill_count}(SF:{stats.sf_wb_count}), "
f"exp={stats.expected_pieces}, arr={stats.arrived_pieces}, "
f"miss={stats.undelivered_pieces}, "
f"miss_wb={stats.undelivered_wb}(full={stats.full_miss}, part={stats.part_miss})"
)
if stats.sf_undelivered:
print(f" SF 差缺: {stats.sf_undelivered} 个运单")
return result
# ============================== Excel 输出 ==============================
# 样式常量(与 compare.py 对齐)
_FONT = "微软雅黑"
_BLUE = "305496"
_HEADER_FILL = PatternFill("solid", fgColor=_BLUE)
_HEADER_FONT = Font(name=_FONT, bold=True, color="FFFFFF", size=11)
_BODY_FONT = Font(name=_FONT, size=10)
_THIN = Side(style="thin", color="D9D9D9")
_BORDER = Border(left=_THIN, right=_THIN, top=_THIN, bottom=_THIN)
def write_result_excel(result: CompareResult, output_path: str | None = None) -> str:
"""将比对结果写入 Excel 文件。
Args:
result: compare_site_date 或 compare_site_batch 的返回值
output_path: 输出路径,为 None 时自动生成:
output/{站}-{日期}-未到数据.xlsx
Returns:
实际写入的文件路径
"""
if output_path is None:
os.makedirs(OUTPUT_DIR, exist_ok=True)
date_tag = result.date.replace(":", "-").replace("batch:", "batch-")
output_path = os.path.join(
OUTPUT_DIR, f"{result.site}-{date_tag}-未到数据.xlsx"
)
wb = Workbook()
ws = wb.active
ws.title = result.site
_write_sheet(ws, result)
wb.save(output_path)
print(f"[db_compare] Excel 已输出: {output_path}")
return output_path
def _write_sheet(ws, result: CompareResult):
"""写单个站点的差缺明细 sheet。"""
s = result.stats
rows = result.rows
# 动态列: 交接单号 | 运单号 | 总件数 | 已到单号1 | 已到单号2 | ...
max_arrived = max((len(r.arrived_list) for r in rows), default=0)
columns = ["交接单号", "运单号", "总件数"] + [
f"已到单号{i + 1}" for i in range(max_arrived)
]
ws.sheet_view.showGridLines = False
# 表头
ws.append(columns)
for c in range(1, len(columns) + 1):
cell = ws.cell(row=1, column=c)
cell.fill = _HEADER_FILL
cell.font = _HEADER_FONT
cell.alignment = Alignment(horizontal="center", vertical="center")
cell.border = _BORDER
# 数据行
for row in rows:
values = {
"交接单号": row.handover_no,
"运单号": row.waybill_no,
"总件数": row.total_pieces,
}
for i, piece in enumerate(row.arrived_list):
values[f"已到单号{i + 1}"] = piece
ws.append([values.get(c, "") for c in columns])
# 格式
for r in range(2, ws.max_row + 1):
for c, col in enumerate(columns, start=1):
cell = ws.cell(row=r, column=c)
cell.font = _BODY_FONT
cell.border = _BORDER
if col == "总件数":
cell.number_format = "#,##0"
cell.alignment = Alignment(horizontal="right", vertical="center")
else:
cell.number_format = "@"
# 列宽
for c, col in enumerate(columns, start=1):
body_lens = [
len(str(ws.cell(row=r, column=c).value or ""))
for r in range(2, ws.max_row + 1)
]
width = min(max([len(str(col))] + body_lens) + 4, 36)
ws.column_dimensions[ws.cell(row=1, column=c).column_letter].width = max(
width, 12
)
ws.freeze_panes = "A2"
# ============================== 全站汇总报表DB 版)=============================
def _stats_to_dict(s: CompareStats) -> dict:
"""CompareStats -> build_summary 要的中文键 stats dict。"""
return {
"运单数": s.waybill_count,
"应到件": s.expected_pieces,
"已到件": s.arrived_pieces,
"未到件": s.undelivered_pieces,
"完全未到": s.full_miss,
"部分未到": s.part_miss,
}
def _target_date_for(site: str) -> str:
"""4 站比对锚点today - actual_offset以实到扫描日为锚与 _site_undelivered_handler 一致)。"""
from inbound_verify import state_store # 懒导入,避免成环
offset = state_store.get_offset(site, "actual")
return (date.today() - timedelta(days=offset)).strftime("%Y-%m-%d")
def _baishi_from_pg(cur, target: str):
"""查百世当日基数baishi_daily_stats+ 当天未到明细undelivered_record 按 ingested_at 过滤)。
返回 (stats_dict_or_None, rows_or_None);基数与明细均无 → (None, None)。
undelivered_record 是 UPSERT 累积表;按 ingested_at::date = target 取当天入库的未到快照
= 当天下载的当前未到,站点已剔除已到),避免累积偏大。
"""
cur.execute(
"SELECT expected_pieces, arrived_pieces, undelivered_pieces "
"FROM baishi_daily_stats WHERE site = %s AND business_date = %s",
("百世", target),
)
basis = cur.fetchone()
cur.execute(
"SELECT waybill_no, piece_no, biz_type, last_scan FROM undelivered_record "
"WHERE site = %s AND ingested_at::date = %s",
("百世", target),
)
detail = cur.fetchall()
if basis is None and not detail:
return (None, None)
exp = basis[0] if basis else None
arr = basis[1] if basis else None
# 未到件优先取基数差baishi_daily_stats.undelivered_pieces与应到/已到同源自洽);
# 基数缺失时退回明细行数。
undel = basis[2] if (basis and basis[2] is not None) else len(detail)
wb_count = len({r[0] for r in detail if r[0]}) # 运单号去重
rows = [
{
"类型": r[2] or "",
"子单号": r[1] or "",
"运单号": r[0] or "",
"最新扫描记录": r[3] or "",
}
for r in detail
]
stats = {
"运单数": wb_count,
"应到件": exp,
"已到件": arr,
"未到件": undel,
"完全未到": None,
"部分未到": None,
}
return (stats, rows)
def build_full_report(date=None) -> str:
"""DB 版全站汇总报表4 站走 DB 比对、百世走 PG复用 compare.build_summary 渲染。
产出 output/应到未到数据.xlsx/report 下载。date=None 时各站按 actual_offset 算锚点(以实到扫描日为锚)。
返回输出路径。"""
from inbound_verify import compare # 复用 build_summary / write_station / OUTFILE
print("[db_compare] 开始生成全站汇总报表 ...")
wb = Workbook()
wb.remove(wb.active)
summary_ws = wb.create_sheet("汇总报表")
results = [] # [(name, stats_dict_or_None)],顺序 ALL_REPORT_SITES
site_targets = {} # name -> target_date汇总表"数据日期"列)
conn = _connect()
cur = conn.cursor()
try:
for name in ALL_REPORT_SITES:
if name == "百世":
target = date or datetime.now().strftime("%Y-%m-%d")
site_targets[name] = target
stats, rows = _baishi_from_pg(cur, target)
results.append((name, stats))
if rows is not None:
compare.write_station(wb.create_sheet(name), BAISHI_COLUMNS, rows)
continue
if name not in SITE_COMPARE_CONFIG:
results.append((name, None))
continue
target = date or _target_date_for(name)
site_targets[name] = target
result = compare_site_date(name, target)
if result is not None:
results.append((name, _stats_to_dict(result.stats)))
_write_sheet(wb.create_sheet(name), result)
else:
results.append((name, None))
finally:
conn.close()
compare.build_summary(
summary_ws,
results,
datetime.now().strftime("%Y-%m-%d %H:%M"),
dates=site_targets,
)
os.makedirs(OUTPUT_DIR, exist_ok=True)
wb.save(compare.OUTFILE)
print(f"[db_compare] 全站汇总已输出: {compare.OUTFILE}")
for name, s in results:
print(f" {name}:未到 {s['未到件']}" if s else f" {name}:无数据,跳过")
return compare.OUTFILE
# ============================== 终端验证入口 ==============================
def main():
"""命令行验证入口:
python -m inbound_verify.db_compare 顺心 2026-07-25
"""
import sys
site = sys.argv[1] if len(sys.argv) > 1 else "顺心"
target_date = sys.argv[2] if len(sys.argv) > 2 else "2026-07-25"
result = compare_site_date(site, target_date)
if result is None:
print(f"{site} {target_date}: 无结果")
return
print(f"\n=== {result.site} {result.date} 差缺明细 ===")
print(f"涉及批次: {result.batches}")
print(f"应到运单: {result.stats.waybill_count} (SF: {result.stats.sf_wb_count})")
print(f"应到件数: {result.stats.expected_pieces}")
print(f"实到件数: {result.stats.arrived_pieces}")
print(f"未到件数: {result.stats.undelivered_pieces}")
print(
f"差缺运单: {result.stats.undelivered_wb} (完全未到: {result.stats.full_miss}, 部分未到: {result.stats.part_miss})"
)
if result.stats.sf_undelivered:
print(f"SF 差缺: {result.stats.sf_undelivered}")
if result.rows:
print(f"\n--- 差缺明细 (共 {len(result.rows)} 条) ---")
for row in result.rows[:20]:
sf = "[SF]" if row.is_sf else ""
arrived_preview = row.arrived_list[:5]
print(
f" {sf} {row.waybill_no}: "
f"应到{row.total_pieces}件, 实到{row.arrived_pieces}"
f" {f'已到: {arrived_preview}' if arrived_preview else ''}"
)
if len(result.rows) > 20:
print(f" ... 还有 {len(result.rows) - 20}")
# 输出 Excel
path = write_result_excel(result)
print(f"\n结果文件: {path}")
if __name__ == "__main__":
main()

92
inbound_verify/domain.py Normal file
View File

@@ -0,0 +1,92 @@
# -*- coding: utf-8 -*-
"""domain.py — 站点 / 文件名 / 列映射的共享配置(单一来源)。
比对compare与入库store都依赖这套配置抽出独立 leaf 模块,
让 store 不必为读配置而依赖整个比对引擎。纯数据,无 state_store / 文件 IO 依赖。
"""
from collections import defaultdict
# 汇总报表覆盖的全部站点4 站在前、百世在末;汇总页图表只取 4 站)
ALL_REPORT_SITES = ["顺心", "中通", "韵达", "安能", "百世"]
# 4 站单站未到明细文件名(百世未到文件由站点直接产出,名为 BAISHI_FILE
SITE_UNDELIVERED_FILE = "{name}-未到数据.xlsx"
BAISHI_FILE = "百世-应到未到货物数据.xlsx"
BAISHI_COLUMNS = ["类型", "子单号", "运单号", "最新扫描记录"]
def arrived_pieces_zhongtong(df):
"""中通实到「运单号」为复合串H + 运单号(12) + 总数(4) + 顺序(4))。
基号 = v[:-8](与应到表运单号对齐),单件 = 整串(每串即一件)。"""
res = defaultdict(set)
for v in df["运单号"]:
v = str(v).strip()
if len(v) > 8 and v[-4:].isdigit():
res[v[:-8]].add(v) # 以完整复合串作为“已到单号”存入
return res
def arrived_pieces_by_cols(wb_col, piece_col):
"""顺心 / 韵达 / 安能:按干净运单列分组,单件 = 子单号 / 扫描单号。
wb_col实到表中与应到运单号对齐的干净列
(顺心=运单号 / 韵达=主单号 / 安能=所属单号)
piece_col实到表中每件货物的单号列子单号 / 扫描单号)"""
def parse(df):
res = defaultdict(set)
for m, s in zip(df[wb_col], df[piece_col]):
m, s = str(m).strip(), str(s).strip()
if m and s:
res[m].add(s)
return res
return parse
STATIONS = [
{
"name": "中通",
"exp": "中通-应到货物数据.xlsx",
"act": "中通-实到货物数据.xlsx",
"exp_qty": "交接件数", # 应到件数口径:交接件数(非录单件数)
"exp_wb": "运单号", # 应到表运单号列(兼作去重键)
"exp_jd": "交接单号", # 未到数据需展示的交接单号
"arrived_pieces": arrived_pieces_zhongtong,
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "顺心",
"exp": "顺心-应到货物数据.xlsx",
"act": "顺心-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("运单号", "子单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "韵达",
"exp": "韵达-应到货物数据.xlsx",
"act": "韵达-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("主单号", "子单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
{
"name": "安能",
"exp": "安能-应到货物数据.xlsx",
"act": "安能-实到货物数据.xlsx",
"exp_qty": "交接件数",
"exp_wb": "运单号",
"exp_jd": "交接单号",
"arrived_pieces": arrived_pieces_by_cols("所属单号", "扫描单号"),
"columns": ["交接单号", "运单号", "总件数"],
},
]
def _site_cfg(name):
"""按名称取 4 站配置(百世不在 STATIONS返回 None"""
return next((c for c in STATIONS if c["name"] == name), None)

View File

@@ -6,7 +6,8 @@
import os
# 项目根目录(以本文件所在位置为基准,与从哪个目录启动脚本无关)
BASE_DIR = os.path.dirname(os.path.abspath(__file__))
# __file__ = <root>/inbound_verify/paths.py → 上两级 = 项目根
BASE_DIR = os.path.dirname(os.path.dirname(os.path.abspath(__file__)))
# 统一的下载 / 输出目录
DOWNLOAD_DIR = os.path.join(BASE_DIR, "downloads")
@@ -17,3 +18,6 @@ CONFIG_PATH = os.path.join(BASE_DIR, "config.yaml")
# 状态存储SQLite阶段0心跳 / 登录态 / 数据态持久化,重启不丢)
STATE_DB_PATH = os.path.join(BASE_DIR, "state", "state.db")
# 错误截图目录(下载流程失败时自动截取,供问题排查)
SCREENSHOT_DIR = os.path.join(BASE_DIR, "logs", "screenshots")

View File

@@ -2,8 +2,8 @@
# 阶段1服务化共享核心。把"启动浏览器 + 各站就绪 + 弹窗 + 心跳初值"
# (launch_and_prepare)、"执行一条任务"(dispatch_task)、"一轮心跳"(run_heartbeat)
# 抽出来,供
# - main_router.py交互菜单模式调试 / 人工操作)
# - server.pyFastAPI 服务模式:常驻 + 接收 API 指令)
# - cli/router.py交互菜单模式调试 / 人工操作)
# - cli/server.pyFastAPI 服务模式:常驻 + 接收 API 指令)
# 共同复用,避免两处重复维护。
#
# 线程模型launch_and_prepare 内 sync_playwright().start() 必须在"持有 Playwright 的
@@ -16,27 +16,23 @@ import socket
import subprocess
import time
import urllib.request
from datetime import datetime, timedelta
from datetime import date, datetime, timedelta
import yaml
from playwright.sync_api import sync_playwright
from paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify.paths import CONFIG_PATH, SCREENSHOT_DIR
import state_store
import site_shunxin
import site_baishi
import site_zto
import site_yunda
import site_anneng
import expected_undelivered # dispatch 的 compare 任务用
from inbound_verify import state_store
from inbound_verify.sites import shunxin, baishi, zto, yunda, anneng
from inbound_verify import compare # dispatch 的 compare 任务用
# 各网页站点首页 URL单一来源取自各站点模块 HOME_URL
SITES_CONFIG = {
"顺心": site_shunxin.HOME_URL,
"百世": site_baishi.HOME_URL,
"中通": site_zto.HOME_URL,
"韵达": site_yunda.HOME_URL,
"顺心": shunxin.HOME_URL,
"百世": baishi.HOME_URL,
"中通": zto.HOME_URL,
"韵达": yunda.HOME_URL,
}
# 站点就绪特征:登录成功进入工作台后的标志性控件
@@ -53,31 +49,6 @@ APP_SITES = {"安能"}
# 心跳间隔(秒)
HEARTBEAT_INTERVAL = 30
# 各站最终数据文件名(探测"数据是否已跑出来");百世为单流程
DATA_FILENAMES = {
"顺心": {
"expected": "顺心-应到货物数据.xlsx",
"actual": "顺心-实到货物数据.xlsx",
"undelivered": "顺心-未到数据.xlsx",
},
"中通": {
"expected": "中通-应到货物数据.xlsx",
"actual": "中通-实到货物数据.xlsx",
"undelivered": "中通-未到数据.xlsx",
},
"韵达": {
"expected": "韵达-应到货物数据.xlsx",
"actual": "韵达-实到货物数据.xlsx",
"undelivered": "韵达-未到数据.xlsx",
},
"安能": {
"expected": "安能-应到货物数据.xlsx",
"actual": "安能-实到货物数据.xlsx",
"undelivered": "安能-未到数据.xlsx",
},
"百世": {"expected": "", "actual": "", "undelivered": "百世-应到未到货物数据.xlsx"},
}
# ============================ 安能启动CDP============================
@@ -105,17 +76,43 @@ def _wait_cdp_up(port, timeout=60.0):
return False
# 【环境兼容】宿主 shellCodex/VS Code 插件、WorkBuddy 等)会向子进程注入一批与业务
# 无关的变量,实测会让安能应用登录后反复弹出“获取试用网点接口报错”:
# - HTTP(S)_PROXY=http://127.0.0.1:8800QuickQ 加速器代理):安能的 wnp.ane56.com
# 接口请求被塞进第三方代理后返回 400/用户未登录;
# - VSCODE_* / CODEX_* / EFC_*VS Code 扩展宿主注入IPC、PID、NLS、ESM 等);
# - NODE_TLS_REJECT_UNAUTHORIZED / DEBUG / RUST_LOG 等宿主调试变量。
# 另ELECTRON_RUN_AS_NODE=1 会把安能当作纯 Node 运行(拒绝 Chromium 参数、启动即
# 退出 rc=9NODE_OPTIONS 同样会干扰。拉起前全部摘掉,尽量还原终端手动启动环境。
_ANNENG_STRIP_PREFIXES = ("VSCODE_", "CODEX_", "EFC_")
_ANNENG_STRIP_EXACT = {
"NODE_OPTIONS",
"ELECTRON_RUN_AS_NODE",
"HTTP_PROXY",
"HTTPS_PROXY",
"ALL_PROXY",
"NO_PROXY",
"NODE_TLS_REJECT_UNAUTHORIZED",
"NODEFAULTCURRENTDIRECTORYINEXEPATH",
"DEBUG",
"RUST_LOG",
"APPLICATION_INSIGHTS_NO_STATSBEAT",
}
def launch_anneng(app_path):
"""以调试模式启动安能 Electron 应用(自动选取空闲端口),返回子进程对象。"""
# 【环境兼容】WorkBuddy 等 shell 会注入 NODE_OPTIONS含 --use-system-ca
# Electron 内置 Node 拒绝该 flag 导致安能启动即退出rc=9
# 拉起前从子进程环境里摘掉 NODE_OPTIONS。
anneng_env = os.environ.copy()
anneng_env.pop("NODE_OPTIONS", None)
anneng_env = {
key: value
for key, value in os.environ.items()
if key not in _ANNENG_STRIP_EXACT and not key.startswith(_ANNENG_STRIP_PREFIXES)
}
port = _find_free_port()
print(f">> 以调试模式启动【安能】应用(端口 {port}{app_path}")
proc = subprocess.Popen([app_path, f"--remote-debugging-port={port}"], env=anneng_env)
site_anneng.set_cdp_port(port)
proc = subprocess.Popen(
[app_path, f"--remote-debugging-port={port}"], env=anneng_env
)
anneng.set_cdp_port(port)
if not _wait_cdp_up(port):
raise RuntimeError(
f"安能应用调试端口 {port} 未就绪——可能应用已在运行(单实例),"
@@ -136,7 +133,7 @@ def probe_site_login(site_name, pages_map):
if site_name not in pages_map:
return False
if site_name == "安能":
return site_anneng.anneng_ready()
return anneng.anneng_ready()
if site_name == "顺心":
return all(
pg.locator(READY_SELECTORS["顺心"]).is_visible(timeout=500)
@@ -151,19 +148,6 @@ def probe_site_login(site_name, pages_map):
return False
def probe_data_file(site_name, kind):
"""探测单站应到/实到数据文件是否存在且为今天。返回 (is_today, mtime_str)。"""
fname = DATA_FILENAMES.get(site_name, {}).get(kind, "")
if not fname:
return (False, "")
path = os.path.join(DOWNLOAD_DIR, fname)
if not os.path.exists(path):
return (False, "")
dt = datetime.fromtimestamp(os.path.getmtime(path))
is_today = dt.date() == datetime.now().date()
return (is_today, dt.strftime("%Y-%m-%d %H:%M:%S"))
# ============================ 运行上下文 ============================
@@ -180,6 +164,7 @@ class RuntimeContext:
sites_to_watch,
debug_mode,
debug_target,
foreground=True,
):
self.pw = pw
self.browser = browser
@@ -189,6 +174,8 @@ class RuntimeContext:
self.sites_to_watch = sites_to_watch
self.debug_mode = debug_mode
self.debug_target = debug_target
# True=任务执行时把 page 置顶交互调试False=后台静默不置顶(服务模式,避免抢焦点)
self.foreground = foreground
def stop(self):
"""关闭 browser + 安能 + Playwright。退出时调用。"""
@@ -237,11 +224,15 @@ def seed_legacy_config():
print(f">> [seed] 从 config.yaml 灌入站点配置: {', '.join(seeded)}")
def launch_and_prepare(debug_mode=False, debug_target=""):
def launch_and_prepare(debug_mode=False, debug_target="", foreground=True):
"""启动 Playwright + 各站 page + 就绪轮询 + 弹窗清理 + 心跳初值,返回 RuntimeContext。
必须在"持有 Playwright 的线程"调用交互模式主线程 / 服务模式 worker 线程
阻塞至所有站点登录就绪才返回
foregroundTrue=任务执行时把 page 置顶交互调试False=后台静默不置顶服务模式
避免抢用户焦点仅控制任务执行阶段的 bring_to_front启动登录/初始弹窗清理的置顶
不受影响启动时窗口需对用户可见以便登录
"""
# 0. 状态库建表/迁移 + 从 config.yaml 灌入站点配置(须在 reset_login_states 等之前)
state_store.init_db()
@@ -295,26 +286,53 @@ def launch_and_prepare(debug_mode=False, debug_target=""):
_launch_args = ["--remote-debugging-port=9223"] if debug_mode else []
browser = pw.chromium.launch(headless=False, args=_launch_args)
context = browser.new_context(viewport={"width": 1920, "height": 1080})
# 默认禁用麦克风/摄像头:在每个页面/iframe 加载前覆盖 getUserMedia 为“直接拒绝”,
# 这样站点(如韵达登录/工作台会请求麦克风)调用时立即 NotAllowedErrorChromium 不再
# 弹出系统授权窗,且麦克风被真正挡住(不是授权给它)。物流工作台无需音视频采集。
context.add_init_script(
"(()=>{const d=()=>Promise.reject(new DOMException('Permission disabled','NotAllowedError'));"
"if(navigator.mediaDevices)navigator.mediaDevices.getUserMedia=d;"
"for(const k of ['getUserMedia','webkitGetUserMedia','mozGetUserMedia']){"
"if(typeof navigator[k]==='function')navigator[k]=function(){return d();};}})();"
)
pages_map = {}
print("\n====================================================")
print("【启动】正在打开各站点页面...")
print("====================================================")
def _open_page(label, attempts=3):
"""开一个页面并 goto(url, domcontentloaded);容忍瞬时 DNS/超时抖动重试,全失败才抛。
domcontentloadedDOM 就绪即返回不等慢资源(广告/图片) load 事件
避免某站 load 30s / 瞬时 DNS 失败导致整个 worker 启动失败就绪轮询自会判登录态"""
last = None
for i in range(1, attempts + 1):
pg = context.new_page()
try:
pg.goto(url, wait_until="domcontentloaded")
return pg
except Exception as e:
last = e
try:
pg.close()
except Exception:
pass
print(f"⚠️ 打开【{label}】第 {i}/{attempts} 次失败: {e}")
if i < attempts:
time.sleep(2)
raise RuntimeError(f"打开【{label}】连续 {attempts} 次失败: {last}")
for site_name, url in active_sites.items():
if site_name == "顺心":
# 顺心:两个归属地账号在同一窗口各开一个标签页
sx_pages = []
for acct in range(1, 3):
print(f">> 正在启动【顺心】账号{acct}标签页: {url}")
sx_page = context.new_page()
sx_page.goto(url)
sx_pages.append(sx_page)
sx_pages.append(_open_page(f"顺心账号{acct}"))
pages_map["顺心"] = sx_pages
else:
print(f">> 正在启动【{site_name}】页面: {url}")
page = context.new_page()
page.goto(url)
pages_map[site_name] = page
pages_map[site_name] = _open_page(site_name)
# 4. 安能 Electron
anneng_proc = None
@@ -333,7 +351,7 @@ def launch_and_prepare(debug_mode=False, debug_target=""):
if "韵达" in pages_map:
try:
pages_map["韵达"].bring_to_front()
site_yunda.yunda_login(pages_map["韵达"])
yunda.yunda_login(pages_map["韵达"])
except Exception as e:
print(f" ⚠️ 韵达前置自动登录模块发生波动: {e}")
@@ -380,6 +398,7 @@ def launch_and_prepare(debug_mode=False, debug_target=""):
sites_to_watch,
debug_mode,
debug_target,
foreground,
)
@@ -403,9 +422,9 @@ def _dismiss_initial_popups(pages_map):
bs_page = pages_map["百世"]
bs_page.bring_to_front()
print(">> 正在处理【百世】初始弹窗(阅读消息 / 配置检查 / 优惠券广告)...")
# 委托给 site_baishi 的专用清理:优惠券广告是全屏居中 modal
# 委托给 baishi 的专用清理:优惠券广告是全屏居中 modal
# 关闭键为 .ant-modal-close纯图标无文字必须点它才能真正关掉。
site_baishi.dismiss_baishi_popups(bs_page)
baishi.dismiss_baishi_popups(bs_page)
print(" ✅ 【百世】初始弹窗处理完成。")
except Exception:
pass
@@ -414,7 +433,7 @@ def _dismiss_initial_popups(pages_map):
yd_page = pages_map["韵达"]
yd_page.bring_to_front()
print(">> 正在检查【韵达】音频设备授权提示...")
site_yunda.dismiss_audio_prompt(yd_page)
yunda.dismiss_audio_prompt(yd_page)
except Exception:
pass
@@ -423,96 +442,276 @@ def _dismiss_initial_popups(pages_map):
def _web_handler(site, download_func):
"""构造网页站任务 handler:单 page 先 bring_to_front 再 download顺心(list) 直接传。"""
"""构造网页站任务 handler
def handler(ctx):
foregroundctx控制任务执行时是否把 page 置顶服务模式后台跑不置顶避免抢用户
焦点交互模式置顶便于调试顺心是 page 列表置顶标志透传给 shunxin_download
"""
def handler(ctx, force=False, date=None):
pg = ctx.pages_map[site]
if not isinstance(pg, list):
if isinstance(pg, list):
# 顺心双账号:置顶与否交给 shunxin_download 在逐账号循环里按 foreground 决定
return download_func(pg, foreground=ctx.foreground, force=force, date=date)
if ctx.foreground:
pg.bring_to_front()
return download_func(pg)
return download_func(pg, force=force, date=date)
return handler
def _site_undelivered_handler(site):
"""4 站未到:下应到+实到 → 比对写 downloads/<站>-未到数据.xlsx。
任一下载失败 清掉旧未到文件返回 False前端不展示陈旧未到"""
"""4 站未到:下应到+实到 → DB 比对 → 写 output/<站>-<日期>-未到数据.xlsx。
应到全量去重已落库则跳过导出因此比对不依赖 Excel 文件走数据库查询
下载成功则返回 True比对失败不影响任务判定数据已入库"""
def handler(ctx):
# 各站下载入口约定返回 True/False顺心历史返回 None视为成功与 dispatch 一致)
exp_ok = TASK_HANDLERS[(site, "expected")](ctx) is not False
def handler(ctx, force=False, date=None):
exp_ok = TASK_HANDLERS[(site, "expected")](ctx, force, date) is not False
act_ok = (
(TASK_HANDLERS[(site, "actual")](ctx) is not False) if exp_ok else False
(TASK_HANDLERS[(site, "actual")](ctx, force, date) is not False)
if exp_ok
else False
)
if exp_ok and act_ok:
return expected_undelivered.write_site_file(site)
stale = os.path.join(
DOWNLOAD_DIR, expected_undelivered.SITE_UNDELIVERED_FILE.format(name=site)
)
if os.path.exists(stale):
os.remove(stale)
if not exp_ok or not act_ok:
return False
# ── 先入库再比对(修复时序:比对须读到本次下载的数据,
# 否则首次/force 时 PG 无当天数据,比对返回 None、不产出 Excel──
try:
_record_business_date(site, "undelivered", date)
except Exception:
pass
try:
from inbound_verify import store # 懒导入,避免成环
if store.ingest_enabled():
store.ingest_task(
site, "undelivered"
) # 4 站 = ingest expected + actual
print(f">> [入库] {site} 前置入库完成")
except Exception as e:
print(f">> [入库] {site} 前置入库失败(不影响比对尝试): {e}")
# ── DB 比对(替代旧 Excel 比对)──
try:
from inbound_verify import db_compare # 懒导入,避免成环
if date:
target_date = date
else:
offset = state_store.get_offset(site, "actual")
target_date = (datetime.now().date() - timedelta(days=offset)).strftime(
"%Y-%m-%d"
)
result = db_compare.compare_site_date(site, target_date)
if result is not None:
db_compare.write_result_excel(result)
else:
print(f">> [未到] {site} {target_date}: 当天无实到数据,跳过比对")
except Exception as e:
print(f">> [未到] {site} DB 比对异常(不影响下载结果): {e}")
return True # 下载成功即返回 True比对失败不影响任务判定
return handler
# 「跑比对」= 纯离线比对(用 downloads/ 现有文件生成全站汇总;下载交由各站定时/手动)。
# 「跑比对」= DB 版全站汇总报表(替代旧 compare.main Excel 路径;下载交由各站定时/手动)。
def _run_db_full_report(date=None):
"""生成 DB 版全站汇总报表output/应到未到数据.xlsx
懒导入 db_comparebest-effort失败只告警返回 True与旧 lambda 契约一致"""
try:
from inbound_verify import db_compare
db_compare.build_full_report(date)
except Exception as e:
print(f">> [跑比对] DB 汇总报表生成失败: {e}")
return True
TASK_HANDLERS = {
("顺心", "expected"): _web_handler("顺心", site_shunxin.shunxin_expected_download),
("顺心", "actual"): _web_handler("顺心", site_shunxin.shunxin_actual_download),
("顺心", "expected"): _web_handler("顺心", shunxin.shunxin_expected_download),
("顺心", "actual"): _web_handler("顺心", shunxin.shunxin_actual_download),
("顺心", "undelivered"): _site_undelivered_handler("顺心"),
("百世", "undelivered"): _web_handler(
"百世", site_baishi.baishi_download_undelivered_data
"百世", baishi.baishi_download_undelivered_data
),
("中通", "expected"): _web_handler("中通", site_zto.zto_expected_download),
("中通", "actual"): _web_handler("中通", site_zto.zto_actual_download),
("中通", "expected"): _web_handler("中通", zto.zto_expected_download),
("中通", "actual"): _web_handler("中通", zto.zto_actual_download),
("中通", "undelivered"): _site_undelivered_handler("中通"),
("韵达", "expected"): _web_handler("韵达", site_yunda.yunda_expected_download),
("韵达", "actual"): _web_handler("韵达", site_yunda.yunda_actual_download),
("韵达", "expected"): _web_handler("韵达", yunda.yunda_expected_download),
("韵达", "actual"): _web_handler("韵达", yunda.yunda_actual_download),
("韵达", "undelivered"): _site_undelivered_handler("韵达"),
("安能", "expected"): lambda ctx: site_anneng.anneng_expected_download(),
("安能", "actual"): lambda ctx: site_anneng.anneng_actual_download(),
(
"安能",
"expected",
): lambda ctx, force=False, date=None: anneng.anneng_expected_download(
force=force, date=date
),
(
"安能",
"actual",
): lambda ctx, force=False, date=None: anneng.anneng_actual_download(
force=force, date=date
),
("安能", "undelivered"): _site_undelivered_handler("安能"),
("__compare__", "compare"): lambda ctx: (expected_undelivered.main() or True),
("__compare__", "compare"): lambda ctx, force=False, date=None: _run_db_full_report(
date
),
}
def _record_business_date(site, kind):
def _record_business_date(site, kind, date=None):
"""下载成功后,把本次数据的业务日期快照写进状态库(供前端/报告显示「是哪天的数据」)。
业务日期 = 下载当天 该数据对应的日期偏移__compare__ 无数据概念跳过
date date否则 = 下载当天 该数据对应的日期偏移__compare__ 无数据概念跳过
kind 写入
expected/actual各写自己一列偏移各取其列
undelivered百世直供 0 undelivered4 站未到由 _site_undelivered_handler
expected/actual各写自己一列
undelivered百世直供当天 undelivered4 站未到由 _site_undelivered_handler
内部连带下了 expected+actual不经 dispatch无业务日期写入故此处一并补写
expected/actual/undelivered 三列actual actual 偏移未到跟随 expected 偏移
顺带置 ready=True让前端不必等心跳即可反映下载成功写入失败仅告警不影响任务判定"""
只写业务日期ready 语义已移交入库成功_persist_to_db 置位此处不再碰 ready"""
if site == "__compare__":
return
today = datetime.now().date()
now = datetime.now().strftime("%Y-%m-%d %H:%M:%S")
def _write(k, off):
biz = (today - timedelta(days=off)).strftime("%Y-%m-%d")
try:
state_store.set_data_state(
site, k, ready=True, generated_at=now, business_date=biz
def _write(k, biz_or_off):
# biz_or_off: int=偏移todayoffstr=已确定业务日期date
biz = (
(today - timedelta(days=biz_or_off)).strftime("%Y-%m-%d")
if isinstance(biz_or_off, int)
else biz_or_off
)
try:
state_store.set_business_date(site, k, biz)
except Exception as e:
print(f">> [状态] 写业务日期失败 {site}/{k}: {e}")
def off(kind_key):
return state_store.get_offset(site, kind_key)
if kind == "expected":
_write("expected", state_store.get_offset(site, "expected"))
_write("expected", date if date else off("expected"))
elif kind == "actual":
_write("actual", state_store.get_offset(site, "actual"))
_write("actual", date if date else off("actual"))
elif site == "百世":
_write("undelivered", 0)
else: # 4 站 undelivered连带补写 expected/actual/undelivered 三列
_write("expected", state_store.get_offset(site, "expected"))
_write("actual", state_store.get_offset(site, "actual"))
_write("undelivered", state_store.get_offset(site, "expected"))
_write("expected", date if date else off("expected"))
_write("actual", date if date else off("actual"))
_write("undelivered", date if date else off("expected"))
def _ready_flags(site):
"""从 PG 业务表派生单站三就绪态ready = DB 数据真相)。
expected/actual = PG 中存在对应 target_datetoday offset的数据
百世 undelivered = baishi_daily_stats 中存在 target_date 的数据
4 undelivered = expected_ready actual_ready派生
PG 不可达时返回全 False降级安全不阻塞心跳
返回 (flags: {kind: bool}, dates: {kind: target_date_str})
dates flags 同源ready=True business_date 即该 target_date
彻底消除 ready business_date 不同源导致的日期标签漂移"""
from inbound_verify import store # 懒导入:避免模块级循环
today = date.today()
today_str = today.isoformat()
if site == "百世":
has_und, _ = store.has_data(site, "undelivered", today_str)
return (
{"expected": False, "actual": False, "undelivered": has_und},
{"undelivered": today_str},
)
exp_off = state_store.get_offset(site, "expected")
act_off = state_store.get_offset(site, "actual")
exp_date = (today - timedelta(days=exp_off)).isoformat()
act_date = (today - timedelta(days=act_off)).isoformat()
has_exp, _ = store.has_data(site, "expected", exp_date)
has_act, _ = store.has_data(site, "actual", act_date)
return (
{"expected": has_exp, "actual": has_act, "undelivered": has_exp and has_act},
{"expected": exp_date, "actual": act_date, "undelivered": exp_date},
)
def _apply_ready(site, flags, dates=None):
"""写入单站就绪态 + 业务日期同源ready 与 business_date 均据 PG + offset 派生)。
ready=True 时同步写入 target_date 作为 business_date消除不同源导致的日期标签漂移
失败仅告警"""
for k, rdy in flags.items():
try:
state_store.set_ready(site, k, rdy)
if rdy and dates and dates.get(k):
state_store.set_business_date(site, k, dates[k])
except Exception as e:
print(f">> [状态] 置就绪态失败 {site}/{k}: {e}")
def capture_error_screenshot(page, site, kind, attempt, error):
"""流程失败时截取当前页面,保存到 logs/screenshots/。
page: Playwright Page 对象安能传 None CDP 分支调用方自行处理
截图失败绝不外抛只打告警不干扰任务重试/清场流程"""
try:
os.makedirs(SCREENSHOT_DIR, exist_ok=True)
ts = datetime.now().strftime("%Y%m%d_%H%M%S")
err_short = (error or "unknown")[:40].replace("/", "_").replace("\\", "_")
fname = f"{site}_{kind}_{ts}_attempt{attempt}_{err_short}.png"
path = os.path.join(SCREENSHOT_DIR, fname)
page.screenshot(path=path, full_page=False)
print(f"📸 【{site}-{kind}】错误截图已保存: {path}")
except Exception as se:
print(f"📸 【{site}-{kind}】截图失败(不影响任务): {se}")
def _refresh_ready(site):
"""入库后立即据 PG 派生并写入该站就绪态(省 30s 心跳等待,与心跳同源)。"""
flags, dates = _ready_flags(site)
_apply_ready(site, flags, dates)
def _persist_to_db(site, kind):
"""下载成功后把本次数据入库 PostgreSQL尽力而为绝不外抛不影响任务判定
- __compare__ 无源数据跳过
- auto_ingest=false 时跳过 PG/cpolar 的开发机
- 懒导入 store 以回避 import 顺序storecompare runtimecompare 共存
- 结果写 state_store.ingest_state /api/status 反映入库健康
所有写库/写状态都包 try/except失败仅告警绝不改变 dispatch_task SUCCESS 判定"""
if site == "__compare__":
return
try:
from inbound_verify import store # 懒导入:冷路径(每下载一次),回避成环
except Exception as e:
print(f">> [warn] 入库模块不可用: {e}")
return
try:
if not store.ingest_enabled(): # 移入 tryconfig.yaml 缺失/损坏时也不外抛
print(">> [入库] 已关闭 (auto_ingest=false),跳过")
return
count = store.ingest_task(site, kind)
# 4 站 undelivered 连带入了 expected+actual按实际入库的类补记 ingest_state
# 否则心跳派生 readyexpected ∧ actual → undelivered会读到陈旧值。
logged = (
["expected", "actual", "undelivered"]
if kind == "undelivered" and site != "百世"
else [kind]
)
for k in logged:
state_store.set_ingest_state(site, k, ok=True, count=count)
_refresh_ready(site) # 入库成功 → 立即据 ingest_state 派生就绪态(与心跳同源)
print(f">> [入库] {site}/{kind} 成功,{count}")
except Exception as e:
print(f">> [warn] 入库失败 {site}/{kind}: {e}")
try:
state_store.set_ingest_state(site, kind, ok=False, error=str(e))
except Exception as e2:
print(f">> [warn] 写入库状态也失败: {e2}")
def dispatch_task(ctx, task_spec):
@@ -535,10 +734,11 @@ def dispatch_task(ctx, task_spec):
if handler is None:
return (state_store.TASK_FAILED, f"未知任务: {site}/{kind}")
try:
ret = handler(ctx)
ret = handler(ctx, bool(task_spec.get("force", False)), task_spec.get("date"))
if ret is False:
return (state_store.TASK_FAILED, "任务执行失败(重试耗尽)")
_record_business_date(site, kind)
_record_business_date(site, kind, task_spec.get("date"))
_persist_to_db(site, kind)
return (state_store.TASK_SUCCESS, None)
except Exception as e:
return (state_store.TASK_FAILED, str(e))
@@ -548,8 +748,11 @@ def dispatch_task(ctx, task_spec):
def run_heartbeat(ctx):
"""一轮心跳:探测各站登录态 + 数据文件,写状态库;登录态变化时提示。
"""一轮心跳:探测各站登录态 + 据 PG 业务表派生数据就绪态;登录态变化时提示。
ready 直接查询 PG 业务表expected_record / actual_record / baishi_daily_stats
目标业务日期是否有数据为唯一依据彻底消除 ingest_state 日期比对带来的每日零点重置
_refresh_ready 在入库瞬间即据 PG 派生 30s 等待心跳同源复核
只在 Playwright 所属线程调用
"""
prev = state_store.get_all_status()
@@ -560,6 +763,5 @@ def run_heartbeat(ctx):
now_login = state_store.LOGIN_IN if logged_in else state_store.LOGIN_OUT
if prev_login and prev_login not in (now_login, state_store.LOGIN_UNKNOWN):
print(f"\n ⚠️【{site_name}】登录态变化: {prev_login}{now_login}")
for kind in ("expected", "actual", "undelivered"):
ready, gen_at = probe_data_file(site_name, kind)
state_store.set_data_state(site_name, kind, ready, gen_at)
flags, dates = _ready_flags(site_name)
_apply_ready(site_name, flags, dates)

View File

@@ -0,0 +1 @@
# inbound_verify.sites — 各承运商站点驱动

View File

@@ -1,13 +1,13 @@
# site_anneng.py
# sites/anneng.py
#
# 安能全网门户Electron 应用)—— 应到货物数据下载。
#
# 与其他站点(网页、由 main_router 用 Playwright 驱动)不同,安能是一个 Electron
# 桌面应用:由 main_router 以调试模式启动(动态空闲端口,经 set_cdp_port 告知本模块),
# 与其他站点(网页、由 runtime 用 Playwright 驱动)不同,安能是一个 Electron
# 桌面应用:由 runtime 以调试模式启动(动态空闲端口,经 set_cdp_port 告知本模块),
# 本模块通过该端口的 CDP 驱动它;「进站交接单查询」「导出下载」等右侧 tab 是
# **独立 webContents**(远程网页),在 /json 里是独立目标。
# 也可脱离 main_router 独立运行(此时用默认端口 9222需已自行启动应用
# .venv/Scripts/python.exe site_anneng.py
# 也可脱离 runtime 独立运行(此时用默认端口 9222需已自行启动应用
# python -m inbound_verify.sites.anneng expected # 或 actual
#
# 应到 = 运单信息导出(每张交接单下应到的运单明细)。整体流程:
# 主页(重庆鱼洞镇) → 运营管理 → 进站管理 → 进站交接单查询(tab)
@@ -39,14 +39,41 @@ import yaml
# Windows 控制台默认 GBK打印中文/emoji 会崩,强制 UTF-8。
sys.stdout.reconfigure(encoding="utf-8")
from paths import DOWNLOAD_DIR, CONFIG_PATH
import state_store
from inbound_verify.paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify import state_store
def with_retry(site_name, label, flow, reset, max_attempts=3):
def _capture_error_screenshot(site, kind, attempt, error):
"""安能 CDP 错误截图best-effort失败仅告警绝不外抛"""
try:
import base64, os
from datetime import datetime
from inbound_verify.paths import SCREENSHOT_DIR
os.makedirs(SCREENSHOT_DIR, exist_ok=True)
pages = list_pages()
if not pages:
return
cdp = CDP(pages[0]["webSocketDebuggerUrl"])
ts = datetime.now().strftime("%Y%m%d_%H%M%S")
err_short = (error or "unknown")[:40].replace("/", "_").replace("\\", "_")
fname = f"{site}_{kind}_{ts}_attempt{attempt}_{err_short}.png"
path = os.path.join(SCREENSHOT_DIR, fname)
result = cdp.call("Page.captureScreenshot", format="png")
with open(path, "wb") as f:
f.write(base64.b64decode(result["data"]))
cdp.close()
print(f"📸 【{site}-{kind}】错误截图已保存: {path}")
except Exception as se:
print(f"📸 【{site}-{kind}】截图失败(不影响任务): {se}")
def with_retry(site_name, label, flow, reset, max_attempts=3, page=None):
"""异常兜底flow 失败 → 重置回初始态 → 重试,最多 max_attempts 次(含首次)。
每次失败都重置含最终放弃那一次既是重试前的清场也保证最终放弃时环境干净
最后一次失败时重置前自动截图保存到 logs/screenshots/供问题排查
安能通过 CDP 截图page 参数忽略保留为统一签名兼容
flow 为零参可调用返回 False 视为失败其余视为成功
返回 True=最终成功False=重试耗尽放弃供调度层判断任务成败
"""
@@ -60,6 +87,11 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return True
except Exception as e:
print(f"⚠️ 【{site_name}-{label}】第 {attempt}/{max_attempts} 次失败: {e}")
if attempt == max_attempts:
try:
_capture_error_screenshot(site_name, label, attempt, str(e))
except Exception:
pass
print(f" → 重置【{site_name}】到初始态,清理环境 ...")
try:
reset()
@@ -73,13 +105,13 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return False
CDP_PORT = 9222 # 默认端口(独立运行 site_anneng.py 时用main_router 启动时会用 set_cdp_port 覆盖
CDP_PORT = 9222 # 默认端口(独立运行python -m inbound_verify.sites.anneng时用runtime 启动时会用 set_cdp_port 覆盖
POLL_INTERVAL = 0.5
DEFAULT_TIMEOUT = 25.0
def set_cdp_port(port):
"""main_router 在启动安能应用后调用,告知本模块实际使用的调试端口。"""
"""runtime 在启动安能应用后调用,告知本模块实际使用的调试端口。"""
global CDP_PORT
CDP_PORT = int(port)
@@ -112,7 +144,10 @@ class CDP:
"""绑定到单个页面目标的同步 CDP 客户端。"""
def __init__(self, ws_url):
self.ws = websocket.create_connection(ws_url)
# socket 级超时Electron 业务 tab 偶发不回包时recv 最多卡 15s 即抛
# WebSocketTimeoutException让上层 wait_until/with_retry 能失败→重试,
# 而不是无限阻塞(曾导致安能下载卡死 ~22 分钟、轮询 300s 截止也无法触发)。
self.ws = websocket.create_connection(ws_url, timeout=15)
self._id = 0
self.call("Runtime.enable")
@@ -186,7 +221,7 @@ def find_main_page_cdp():
def anneng_ready():
"""安能主页是否就绪(供 main_router 的就绪轮询调用)。
"""安能主页是否就绪(供 runtime 的就绪轮询调用)。
判据能连上 CDP 且主页出现站点名称控件 title+fontsizenum div
连不上或未就绪一律返回 False不抛异常就绪轮询会反复调用
@@ -852,13 +887,18 @@ def _load_query_days():
return max(1, days)
def anneng_expected_download():
def anneng_expected_download(force=False, date=None):
"""安能:应到货物数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry("安能", "应到", anneng_expected_download_impl, anneng_reset)
return with_retry(
"安能",
"应到",
lambda: anneng_expected_download_impl(force=force, date=date),
anneng_reset,
)
def anneng_expected_download_impl():
def anneng_expected_download_impl(force=False, date=None):
"""安能:应到货物数据(运单信息)下载,完整流程(单次执行,无重试;供自动化测试用)。"""
print("\n▶ 开始执行【安能 - 应到货物数据下载】任务 ...")
download_dir = DOWNLOAD_DIR
@@ -869,14 +909,32 @@ def anneng_expected_download_impl():
offset = state_store.get_offset("安能")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_str = f"{target.year}-{target.month:02d}-{target.day:02d}"
start_str = target_str
today_str = target_str
print(f">> 查询日期: [{target_str}]偏移 {offset}0=今天")
src = f"指定 {date}" if date else f"偏移 {offset}0=今天"
print(f">> 查询日期: [{target_str}]{src}")
main_cdp = find_main_page_cdp()
export_times = []
# 【去重】加载本站已落库交接单号force=True 或查询失败时 existing=空集(不去重)
if force:
existing = set()
print(">> [去重] 强制重下,跳过去重。")
else:
try:
from inbound_verify import store
existing = store.get_existing_handover_nos("安能")
except Exception as _e:
existing = set()
print(f">> [去重] 加载已落库交接单号失败,本次不去重: {_e}")
try:
# 1) 确认主页就绪并导航到“进站交接单查询”
wait_home_ready(main_cdp)
@@ -915,6 +973,10 @@ def anneng_expected_download_impl():
for i, ewbs_no in enumerate(target_ids, start=1):
print(f" ⏳ [{i}/{len(target_ids)}] 交接单号 {ewbs_no}")
# 【去重】已落库则跳过:不双击、不导出、不 append export_times
if ewbs_no in existing:
print(f" ⏭️ 交接单号 {ewbs_no} 已落库,跳过。")
continue
activate_tab(tab_cdp, "交接单信息")
time.sleep(0.3)
if not dblclick_jiaojie_dan_row(tab_cdp, ewbs_no):
@@ -933,6 +995,11 @@ def anneng_expected_download_impl():
close_tab_by_label(main_cdp, "进站交接单查询")
time.sleep(0.8)
# 【去重兜底】全部已落库/无数据 → 无导出任务,查询 tab 已关,跳过下载段
if not export_times:
print(">> 本次无新交接单需导出(全部已落库或无数据),结束。")
return True
# 5) 打开导出下载 tab轮询并下载
print(">> 打开【导出下载】tab ...")
export_cdp = ensure_tab_open(main_cdp, "导出下载", EXPORT_TAB_URL_HINT)
@@ -973,6 +1040,7 @@ def set_scan_date(cdp, placeholder, value, target_ymd=None):
返回 True 表示写入并校验成功否则 False
"""
import re as _re
ph = json.dumps(placeholder)
val_json = json.dumps(value)
if target_ymd is None:
@@ -998,7 +1066,9 @@ def set_scan_date(cdp, placeholder, value, target_ymd=None):
"(() => {const inp=[...document.querySelectorAll('input')]"
f".find(i=>i.placeholder==={ph}); if(!inp) return false; inp.focus();"
"let sel=true; try{ sel=document.execCommand('selectAll'); }catch(e){ try{inp.select();}catch(_){ sel=false; } }"
"let done=false; try{ done=document.execCommand('insertText',false," + val_json + "); }catch(e){ done=false; }"
"let done=false; try{ done=document.execCommand('insertText',false,"
+ val_json
+ "); }catch(e){ done=false; }"
"if(!done){ const s=Object.getOwnPropertyDescriptor(window.HTMLInputElement.prototype,'value').set;"
f" s.call(inp,{val_json}); inp.dispatchEvent(new Event('input',{{bubbles:true}})); }}"
"return inp.value;})()"
@@ -1016,7 +1086,9 @@ def set_scan_date(cdp, placeholder, value, target_ymd=None):
return True
time.sleep(0.4)
# 全部尝试失败:告警,避免静默下成「今天」
print(f" ⚠ set_scan_date 未能将 {placeholder} 设为目标日期 {target_ymd}(请检查 DatePicker 是否就绪)")
print(
f" ⚠ set_scan_date 未能将 {placeholder} 设为目标日期 {target_ymd}(请检查 DatePicker 是否就绪)"
)
return False
@@ -1158,8 +1230,12 @@ def wait_scan_form_ready(cdp, timeout=20.0):
f".find(i=>i.placeholder==={json.dumps(ph)}); if(!inp) return false; inp.focus(); return true;}})()"
)
try:
wait_until(cdp, "!!document.querySelector('.ant-picker-panel')",
"预热面板打开", timeout=4.0)
wait_until(
cdp,
"!!document.querySelector('.ant-picker-panel')",
"预热面板打开",
timeout=4.0,
)
except Exception:
pass
cdp.eval(
@@ -1221,13 +1297,15 @@ def _save_actual(rows, download_dir):
print("====================================================")
def anneng_actual_download():
def anneng_actual_download(force=False, date=None):
"""安能:实到数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry("安能", "实到", anneng_actual_download_impl, anneng_reset)
return with_retry(
"安能", "实到", lambda: anneng_actual_download_impl(date=date), anneng_reset
)
def anneng_actual_download_impl():
def anneng_actual_download_impl(date=None):
"""安能:实到数据(网点到件扫描,子单)下载,完整流程(单次执行,无重试;供自动化测试用)。"""
print("\n▶ 开始执行【安能 - 实到货物数据下载】任务 ...")
download_dir = DOWNLOAD_DIR
@@ -1238,10 +1316,14 @@ def anneng_actual_download_impl():
offset = state_store.get_offset("安能", "actual")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
start_str = f"{target.year}/{target.month:02d}/{target.day:02d} 00:00:00"
end_str = f"{target.year}/{target.month:02d}/{target.day:02d} 23:59:59"
print(f">> 扫描日期: [{start_str}{end_str}]偏移 {offset}0=今天")
src = f"指定 {date}" if date else f"偏移 {offset}0=今天"
print(f">> 扫描日期: [{start_str}{end_str}]{src}")
main_cdp = find_main_page_cdp()
try:

View File

@@ -1,17 +1,18 @@
# site_baishi.py
# sites/baishi.py
import os
import yaml
from playwright.sync_api import sync_playwright
from paths import DOWNLOAD_DIR, CONFIG_PATH
import state_store
from inbound_verify.paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify import state_store
def with_retry(site_name, label, flow, reset, max_attempts=3):
def with_retry(site_name, label, flow, reset, max_attempts=3, page=None):
"""异常兜底flow 失败 → 重置回初始态 → 重试,最多 max_attempts 次(含首次)。
每次失败都重置含最终放弃那一次既是重试前的清场也保证最终放弃时环境干净
最后一次失败时重置前自动截图保存到 logs/screenshots/供问题排查
flow 为零参可调用返回 False 视为失败其余视为成功
返回 True=最终成功False=重试耗尽放弃供调度层判断任务成败
"""
@@ -25,6 +26,13 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return True
except Exception as e:
print(f"⚠️ 【{site_name}-{label}】第 {attempt}/{max_attempts} 次失败: {e}")
if attempt == max_attempts and page is not None:
try:
from inbound_verify.runtime import capture_error_screenshot
capture_error_screenshot(page, site_name, label, attempt, str(e))
except Exception:
pass
print(f" → 重置【{site_name}】到初始态,清理环境 ...")
try:
reset()
@@ -38,7 +46,7 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return False
# 站点首页 URL异常兜底重置用也供 main_router 的 SITES_CONFIG 引用(单一来源)
# 站点首页 URL异常兜底重置用也供 runtime 的 SITES_CONFIG 引用(单一来源)
HOME_URL = "https://v5.800best.com"
@@ -137,14 +145,17 @@ def _close_tab(page, tab_name):
print(f" ⚠️ 关闭标签页【{tab_name}】时出错: {e}")
def baishi_download_undelivered_data(page):
"""百世:一键提取应到未到(当日未扫)数据(内部含异常兜底重试,路由层无感)。"""
def baishi_download_undelivered_data(page, force=False, date=None):
"""百世:一键提取应到未到(当日未扫)数据(内部含异常兜底重试,路由层无感)。
date 形参仅为对齐统一透传签名百世固定下载当天忽略"""
return with_retry(
"百世",
"应到未到",
lambda: baishi_download_undelivered_data_impl(page),
lambda: baishi_reset(page),
page=page,
)
@@ -181,8 +192,12 @@ def baishi_download_undelivered_data_impl(page):
# 10(应扫) 11(已扫) 12(未扫) 13(率)。(与下方 nth(12) 未扫同源)
try:
_first_row = page.locator(".ant-table-tbody > tr").first
_exp_txt = _first_row.locator("td").nth(10).inner_text().strip().replace(",", "")
_arr_txt = _first_row.locator("td").nth(11).inner_text().strip().replace(",", "")
_exp_txt = (
_first_row.locator("td").nth(10).inner_text().strip().replace(",", "")
)
_arr_txt = (
_first_row.locator("td").nth(11).inner_text().strip().replace(",", "")
)
def _to_int(v):
try:
@@ -194,6 +209,11 @@ def baishi_download_undelivered_data_impl(page):
if _exp_n > 0:
state_store.set_setting("百世", "scan_expected_pieces", str(_exp_n))
state_store.set_setting("百世", "scan_arrived_pieces", str(_arr_n))
from inbound_verify import (
store,
) # 直接落库 PG一步不绕 state_store→store
store.upsert_baishi_daily_stats(_exp_n, _arr_n)
print(f" 已记录百世应到/实到基数:应扫 {_exp_n} / 已扫 {_arr_n}")
except Exception as _e:
# 抓取失败绝不影响未到明细下载主流程

View File

@@ -1,4 +1,4 @@
# site_shunxin.py
# sites/shunxin.py
import os
import re
@@ -7,14 +7,15 @@ import yaml
from datetime import datetime, timedelta
import pandas as pd
from paths import DOWNLOAD_DIR, CONFIG_PATH
import state_store
from inbound_verify.paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify import state_store
def with_retry(site_name, label, flow, reset, max_attempts=3):
def with_retry(site_name, label, flow, reset, max_attempts=3, page=None):
"""异常兜底flow 失败 → 重置回初始态 → 重试,最多 max_attempts 次(含首次)。
每次失败都重置含最终放弃那一次既是重试前的清场也保证最终放弃时环境干净
最后一次失败时重置前自动截图保存到 logs/screenshots/供问题排查
flow 为零参可调用返回 False 视为失败其余视为成功
返回 True=最终成功False=重试耗尽放弃供调度层判断任务成败
"""
@@ -28,6 +29,13 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return True
except Exception as e:
print(f"⚠️ 【{site_name}-{label}】第 {attempt}/{max_attempts} 次失败: {e}")
if attempt == max_attempts and page is not None:
try:
from inbound_verify.runtime import capture_error_screenshot
capture_error_screenshot(page, site_name, label, attempt, str(e))
except Exception:
pass
print(f" → 重置【{site_name}】到初始态,清理环境 ...")
try:
reset()
@@ -41,7 +49,7 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return False
# 站点首页 URL异常兜底重置用也供 main_router 的 SITES_CONFIG 引用(单一来源)
# 站点首页 URL异常兜底重置用也供 runtime 的 SITES_CONFIG 引用(单一来源)
HOME_URL = "https://sxne.sxjdfreight.com"
@@ -88,6 +96,64 @@ def _remove_if_exists(path):
pass
def _shunxin_navigate_picker_to_target(page, date_str, max_flips=24):
"""顺心 Ant Design 单月面板跨月导航(车辆点到 / 卸车扫描记录 共用同一组件)。
读面板头部 .ant-picker-year-btn / .ant-picker-month-btn 得当前显示的年月按与目标
年月的差值点 .ant-picker-header-prev-btn上月/ .ant-picker-header-next-btn下月
翻到目标月视窗返回 True 表示当前视窗已是目标月目标格子随后可见可点
调用前提开始/结束时间输入已点开.ant-picker-dropdown:visible 已就绪
"""
try:
ty, tm = (int(x) for x in date_str.split("-")[:2])
except Exception:
return True # 解析不出就不翻,交由后续 cell.click 自行成败
drop = page.locator(".ant-picker-dropdown:visible")
for _ in range(max_flips):
try:
cur_y = int(
re.search(
r"\d+", drop.locator(".ant-picker-year-btn").first.inner_text()
).group()
)
cur_m = int(
re.search(
r"\d+", drop.locator(".ant-picker-month-btn").first.inner_text()
).group()
)
except Exception:
return False
cur = cur_y * 12 + (cur_m - 1)
tgt = ty * 12 + (tm - 1)
if cur == tgt:
return True
btn_sel = (
".ant-picker-header-prev-btn"
if tgt < cur
else ".ant-picker-header-next-btn"
)
drop.locator(btn_sel).first.click()
page.wait_for_timeout(300)
return False
def _shunxin_pick_date(page, date_str):
"""在已打开的顺心 Ant Design 日期浮层上选中指定日期格子(含跨月翻月)。
目标格子不在当前月视窗跨月先调 _shunxin_navigate_picker_to_target 翻到目标月
再点格子同月则直接点与中通 _zto_flip_to_target_month 思路对称适配 Ant Design 面板
"""
cell = page.locator(f".ant-picker-dropdown:visible td[title='{date_str}']").first
if not cell.is_visible():
print(f" 目标日期 {date_str} 不在当前月视窗,正在翻月导航 ...")
if not _shunxin_navigate_picker_to_target(page, date_str):
raise RuntimeError(f"翻月后仍无法定位目标日期格子 {date_str}")
cell = page.locator(
f".ant-picker-dropdown:visible td[title='{date_str}']"
).first
cell.click()
def shunxin_belonging(page):
"""读取顺心当前账号的归属网点名(仅在首页可见,须在导航离开首页前调用)。
@@ -149,7 +215,7 @@ def shunxin_merge_final(kind, tags):
pass
def shunxin_expected_download(pages):
def shunxin_expected_download(pages, foreground=True, force=False, date=None):
"""顺心:应到货物数据下载(双账号/双归属地,内部含异常兜底重试与数据融合)。
pages 为该站点的 page 列表双账号在同一窗口的各一个标签页
@@ -169,13 +235,17 @@ def shunxin_expected_download(pages):
)
for idx, (pg, tag) in enumerate(zip(pages, tags), start=1):
if foreground:
pg.bring_to_front()
print(f"\n========== 顺心 · 账号{idx}{tag})应到数据下载 ==========")
ok = with_retry(
f"顺心-{tag}",
"应到",
lambda p=pg, t=tag: shunxin_expected_download_impl(p, out_tag=t),
lambda p=pg, t=tag, f=force, d=date: shunxin_expected_download_impl(
p, out_tag=t, force=f, date=d
),
lambda p=pg: shunxin_reset(p),
page=pg,
)
if not ok:
return False # 某账号重试耗尽 → 整体失败,不融合(避免部分数据)
@@ -184,7 +254,7 @@ def shunxin_expected_download(pages):
return True
def shunxin_expected_download_impl(page, out_tag=""):
def shunxin_expected_download_impl(page, out_tag="", force=False, date=None):
"""顺心:应到货物数据下载(单次执行,无重试;供自动化测试探测原始结果用)。
out_tag 为归属地标签时合并产物命名为顺心-{out_tag}-应到货物数据.xlsx
@@ -221,27 +291,27 @@ def shunxin_expected_download_impl(page, out_tag=""):
# 2. 读取服务端日期偏移0=今天1=昨天…),单日范围:起止同日
offset = state_store.get_offset("顺心")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_str = target.strftime("%Y-%m-%d")
start_date_str = target_str
today_str = target_str
print(f">> 正在设置查询日期: [{target_str}]偏移 {offset}0=今天...")
src = f"指定 {date}" if date else f"偏移 {offset}0=今天"
print(f">> 正在设置查询日期: [{target_str}]{src}...")
# 分两步精准呼出和点击时间控件
print(" >> 设置起始时间...")
page.get_by_placeholder("开始时间").click()
page.wait_for_timeout(500)
page.locator(
f".ant-picker-dropdown:visible td[title='{start_date_str}']"
).first.click()
_shunxin_pick_date(page, start_date_str)
page.wait_for_timeout(300)
print(" >> 设置截止时间...")
page.get_by_placeholder("结束时间").click()
page.wait_for_timeout(500)
page.locator(
f".ant-picker-dropdown:visible td[title='{today_str}']"
).first.click()
_shunxin_pick_date(page, today_str)
page.wait_for_timeout(300)
# 确认日期
@@ -308,12 +378,44 @@ def shunxin_expected_download_impl(page, out_tag=""):
count = waybill_btns.count()
print(f">> 共发现 {count} 个班次需要导出。")
# 【去重】加载本站已落库交接单号force=True 或查询失败时 existing=空集(不去重)。
# 两账号共享同一集合(班次号/交接单号跨归属地不重叠)。
if force:
existing = set()
print(">> [去重] 强制重下,跳过去重。")
else:
try:
from inbound_verify import store
existing = store.get_existing_handover_nos("顺心")
except Exception as _e:
existing = set()
print(f">> [去重] 加载已落库交接单号失败,本次不去重: {_e}")
for i in range(count):
print(f" ⏳ 正在处理第 {i+1}/{count} 个班次...")
waybill_btns.nth(i).click()
page.locator("label[title='运单查询']").wait_for(state="visible")
# 【方式1】运单列表界面已加载读交接单号RTS 开头)→ 已落库则退回列表跳过。
# 交接单号格式 RTS\d{3}WJ\d+(如 RTS023WJ374837用 [A-Z0-9]+ 连续匹配整段。
# 读不到DOM 变动/未渲染)则 handover_no 为空 → 不跳过(安全降级,继续导出)。
handover_no = ""
try:
_txt = page.locator("text=/RTS\\d+/").first.inner_text(timeout=3000)
_m = re.search(r"RTS[A-Z0-9]+", _txt)
if _m:
handover_no = _m.group(0)
except Exception:
pass
print(f" -> 运单列表交接单号:{handover_no or '(未读到,不去重)'}")
if handover_no and handover_no in existing:
print(f" ⏭️ 交接单号 {handover_no} 已落库,跳过提交导出。")
page.get_by_role("tab", name="车辆点到").click()
page.wait_for_timeout(500)
continue
# 4. 执行导出流程
page.get_by_role("button", name="export 导出").click()
@@ -341,6 +443,11 @@ def shunxin_expected_download_impl(page, out_tag=""):
_close_tab(page, "运单列表")
_close_tab(page, "车辆点到")
# 【去重兜底】全部已落库/无数据 → 无导出任务,标签页已关,跳过下载轮询
if not export_times:
print(">> 本次无新班次需导出(全部已落库或无数据),结束。")
return True
# 6. 前往数据导出页面去下载
print(">> 正在前往【数据导出】界面...")
page.locator("a[href='/dataExport']").click()
@@ -476,7 +583,7 @@ def shunxin_expected_download_impl(page, out_tag=""):
return False
def shunxin_actual_download(pages):
def shunxin_actual_download(pages, foreground=True, force=False, date=None):
"""顺心:实到货物数据下载(双账号/双归属地,内部含异常兜底重试与数据融合)。
shunxin_expected_download 同构读归属地 去重校验 顺序各账号下载
@@ -494,13 +601,17 @@ def shunxin_actual_download(pages):
)
for idx, (pg, tag) in enumerate(zip(pages, tags), start=1):
if foreground:
pg.bring_to_front()
print(f"\n========== 顺心 · 账号{idx}{tag})实到数据下载 ==========")
ok = with_retry(
f"顺心-{tag}",
"实到",
lambda p=pg, t=tag: shunxin_actual_download_impl(p, out_tag=t),
lambda p=pg, t=tag, d=date: shunxin_actual_download_impl(
p, out_tag=t, date=d
),
lambda p=pg: shunxin_reset(p),
page=pg,
)
if not ok:
return False
@@ -509,7 +620,7 @@ def shunxin_actual_download(pages):
return True
def shunxin_actual_download_impl(page, out_tag=""):
def shunxin_actual_download_impl(page, out_tag="", date=None):
"""顺心:实到货物数据下载(单次执行,无重试;供自动化测试探测原始结果用)。
out_tag 为归属地标签时合并产物命名为顺心-{out_tag}-实到货物数据.xlsx
@@ -542,27 +653,27 @@ def shunxin_actual_download_impl(page, out_tag=""):
# 2. 读取服务端日期偏移0=今天1=昨天…),单日范围:起止同日
offset = state_store.get_offset("顺心", "actual")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_str = target.strftime("%Y-%m-%d")
start_date_str = target_str
today_str = target_str
print(f">> 正在设置查询日期: [{target_str}]偏移 {offset}0=今天...")
src = f"指定 {date}" if date else f"偏移 {offset}0=今天"
print(f">> 正在设置查询日期: [{target_str}]{src}...")
# 分两步精准呼出和点击时间控件
print(" >> 设置起始时间...")
page.get_by_placeholder("开始时间").click()
page.wait_for_timeout(500)
page.locator(
f".ant-picker-dropdown:visible td[title='{start_date_str}']"
).first.click()
_shunxin_pick_date(page, start_date_str)
page.wait_for_timeout(300)
print(" >> 设置截止时间...")
page.get_by_placeholder("结束时间").click()
page.wait_for_timeout(500)
page.locator(
f".ant-picker-dropdown:visible td[title='{today_str}']"
).first.click()
_shunxin_pick_date(page, today_str)
page.wait_for_timeout(300)
# 确认日期

View File

@@ -1,4 +1,4 @@
# site_yunda.py
# sites/yunda.py
import os
import re
@@ -7,14 +7,15 @@ import yaml
from datetime import datetime, timedelta
import pandas as pd
from paths import DOWNLOAD_DIR, CONFIG_PATH
import state_store
from inbound_verify.paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify import state_store
def with_retry(site_name, label, flow, reset, max_attempts=3):
def with_retry(site_name, label, flow, reset, max_attempts=3, page=None):
"""异常兜底flow 失败 → 重置回初始态 → 重试,最多 max_attempts 次(含首次)。
每次失败都重置含最终放弃那一次既是重试前的清场也保证最终放弃时环境干净
最后一次失败时重置前自动截图保存到 logs/screenshots/供问题排查
flow 为零参可调用返回 False 视为失败其余视为成功
返回 True=最终成功False=重试耗尽放弃供调度层判断任务成败
"""
@@ -28,6 +29,13 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return True
except Exception as e:
print(f"⚠️ 【{site_name}-{label}】第 {attempt}/{max_attempts} 次失败: {e}")
if attempt == max_attempts and page is not None:
try:
from inbound_verify.runtime import capture_error_screenshot
capture_error_screenshot(page, site_name, label, attempt, str(e))
except Exception:
pass
print(f" → 重置【{site_name}】到初始态,清理环境 ...")
try:
reset()
@@ -41,7 +49,7 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return False
# 站点首页 URL异常兜底重置用也供 main_router 的 SITES_CONFIG 引用(单一来源)
# 站点首页 URL异常兜底重置用也供 runtime 的 SITES_CONFIG 引用(单一来源)
HOME_URL = "https://ky-sso.yunda56.com"
@@ -77,6 +85,93 @@ def _remove_if_exists(path):
pass
def _yunda_pick_laydate_new(ws_frame, page, date_ymd):
"""应到(新版 laydate .layui-laydate在已打开的面板上选中指定日期含跨月翻月
目标格子 td[lay-ymd='YYYY-M-D']非补零不在当前月视窗时 .laydate-set-ym
当前年月形如2026年7月按差值点 .laydate-prev-m / .laydate-next-m 翻到目标月
再点格子同月则直接点调用前提#startTime/#endTime 已点开,.layui-laydate:visible 就绪。
"""
cal = ws_frame.locator(".layui-laydate:visible").first
cell = cal.locator(f"td[lay-ymd='{date_ymd}']").first
if not cell.is_visible():
ty, tm = (int(x) for x in date_ymd.split("-")[:2])
print(f" 目标日期 {date_ymd} 不在当前月视窗,正在翻月导航 ...")
for _ in range(24):
nums = re.findall(r"\d+", cal.locator(".laydate-set-ym").first.inner_text())
if len(nums) >= 2:
cur_y, cur_m = int(nums[0]), int(nums[1])
if cur_y == ty and cur_m == tm:
break
cur = cur_y * 12 + (cur_m - 1)
btn = (
".laydate-prev-m"
if (ty * 12 + (tm - 1)) < cur
else ".laydate-next-m"
)
cal.locator(btn).first.click()
page.wait_for_timeout(300)
cell = cal.locator(f"td[lay-ymd='{date_ymd}']").first
cell.click()
def _yunda_pick_laydate_old(ws_frame, page, date_ymd):
"""实到(旧版 laydate #laydate_box在已打开的面板上选中指定日期含跨月翻月
目标格子 td[y][m][d]非补零不在当前月视窗时 #laydate_y/#laydate_m 输入框值
形如202607得当前年月按差值点 #laydate_MM 内 .laydate_chprev /
.laydate_chnext 翻到目标月再点格子同月则直接点调用前提#startDate/#endDate
已点开force=True#laydate_box:visible 就绪。
"""
box = ws_frame.locator("#laydate_box:visible").first
ty, tm, td = (int(x) for x in date_ymd.split("-")[:3])
cell = box.locator(f"td[y='{ty}'][m='{tm}'][d='{td}']").first
if not cell.is_visible():
print(f" 目标日期 {date_ymd} 不在当前月视窗,正在翻月导航 ...")
for _ in range(24):
yv = box.locator("#laydate_y").first.evaluate("e=>e.value")
mv = box.locator("#laydate_m").first.evaluate("e=>e.value")
cur_y = int(re.search(r"\d+", yv).group())
cur_m = int(re.search(r"\d+", mv).group())
if cur_y == ty and cur_m == tm:
break
cur = cur_y * 12 + (cur_m - 1)
btn = ".laydate_chprev" if (ty * 12 + (tm - 1)) < cur else ".laydate_chnext"
box.locator(f"#laydate_MM {btn}").first.click()
page.wait_for_timeout(300)
cell = box.locator(f"td[y='{ty}'][m='{tm}'][d='{td}']").first
cell.click()
def _resolve_export_frame(ws_frame):
"""定位韵达数据导出面板内嵌的 iframe。
韵达改版后导出面板换用 Element UI全选/向右转移/导出按钮实到流程的面板
iframe 名为 myFrame已验证应到流程历史上为 target1这里短超时轮流探测
返回首个出现全选按钮的 frame都未命中则 dump 面板内所有 iframe 名便于排查
"""
for name in ("myFrame", "target1"):
frame = ws_frame.frame_locator(f'iframe[name="{name}"]')
try:
frame.locator("button.el-button", has_text="全选").first.wait_for(
state="visible", timeout=5000
)
print(f" [导出iframe] 命中 iframe[name={name}]")
return frame
except Exception:
continue
try:
names = ws_frame.locator("iframe").evaluate_all(
"els => els.map(e => e.name || '(无name)')"
)
print(
f" [导出iframe] myFrame/target1 均未命中全选;面板 iframe 名: {names}"
)
except Exception as e:
print(f" [导出iframe] dump 失败: {e}")
return ws_frame.frame_locator('iframe[name="myFrame"]')
def yunda_login(page):
"""韵达自动登录:未登录则填充表单并提交,已登录则跳过。"""
print(">> 正在检查韵达登录状态...")
@@ -140,18 +235,19 @@ def yunda_smart_menu_click(page, menu_path):
page.wait_for_timeout(1000)
def yunda_expected_download(page):
def yunda_expected_download(page, force=False, date=None):
"""韵达:应到货物数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry(
"韵达",
"应到",
lambda: yunda_expected_download_impl(page),
lambda: yunda_expected_download_impl(page, force=force, date=date),
lambda: yunda_reset(page),
page=page,
)
def yunda_expected_download_impl(page):
def yunda_expected_download_impl(page, force=False, date=None):
"""韵达:应到货物数据下载(单次执行,无重试;供自动化测试探测原始结果用)。"""
print("\n▶ 开始执行【韵达 - 应到货物数据下载】任务...")
@@ -182,11 +278,15 @@ def yunda_expected_download_impl(page):
# 2. 读取服务端日期偏移0=今天1=昨天…),单日范围:起止同日
offset = state_store.get_offset("韵达")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
target_ymd = f"{target.year}-{target.month}-{target.day}"
start_date_ymd = target_ymd
today_ymd = target_ymd
print(f">> 设置查询日期: [{target_ymd}]偏移 {offset}0=今天")
src = f"指定 {date}" if date else f"偏移 {offset}0=今天"
print(f">> 设置查询日期: [{target_ymd}]{src}")
# 设定起始时间
print(" >> 设置起始时间...")
@@ -195,7 +295,7 @@ def yunda_expected_download_impl(page):
calendar1 = ws_frame.locator(".layui-laydate:visible").first
calendar1.wait_for(state="visible", timeout=5000)
calendar1.locator(f"td[lay-ymd='{start_date_ymd}']").click()
_yunda_pick_laydate_new(ws_frame, page, start_date_ymd)
calendar1.locator(".laydate-btns-confirm").click()
page.wait_for_timeout(400)
@@ -205,7 +305,7 @@ def yunda_expected_download_impl(page):
calendar2 = ws_frame.locator(".layui-laydate:visible").first
calendar2.wait_for(state="visible", timeout=5000)
calendar2.locator(f"td[lay-ymd='{today_ymd}']").click()
_yunda_pick_laydate_new(ws_frame, page, today_ymd)
calendar2.locator(".laydate-btns-confirm").click()
page.wait_for_timeout(500)
@@ -254,6 +354,19 @@ def yunda_expected_download_impl(page):
row_count = main_rows.count()
print(f">> 当前视窗共捕获到活跃交接单记录: {row_count}")
# 【去重】加载本站已落库交接单号force=True 或查询失败时 existing=空集(不去重)
if force:
existing = set()
print(">> [去重] 强制重下,跳过去重。")
else:
try:
from inbound_verify import store
existing = store.get_existing_handover_nos("韵达")
except Exception as _e:
existing = set()
print(f">> [去重] 加载已落库交接单号失败,本次不去重: {_e}")
# 5. 逐行双击并提交导出
for i in range(row_count):
print(f" ⏳ 正在处理第 {i+1}/{row_count} 个交接单模块...")
@@ -261,6 +374,11 @@ def yunda_expected_download_impl(page):
raw_no = current_row.locator("td").nth(1).inner_text().strip()
# 【去重】已落库的交接单号不再提交导出任务
if raw_no in existing:
print(f" ⏭️ 交接单号 {raw_no} 已落库,跳过提交导出。")
continue
# 跳过已绑定的交接单
bind_status = current_row.locator("td").nth(2).inner_text().strip()
print(f" -> 交接单号: {raw_no} [绑定状态: {bind_status}]")
@@ -274,8 +392,11 @@ def yunda_expected_download_impl(page):
ws_frame.locator("#docSum").wait_for(state="visible", timeout=15000)
page.wait_for_timeout(500)
# 导出弹窗双层重试:外层重新打开面板,内层重新提交
# 区分字段漏选(补点全选)与字段列表消失(重新打开面板)。
# 导出弹窗重试:外层重新打开面板(最多 3 次)
# 韵达站点已将导出面板从 jQuery(.allRight/#submitbutton) 改版为 Element UI
# 与实到流程同一导出组件iframe=myFrame
# 全选(button“全选”) → 向右转移(i.el-icon-d-arrow-right) → 导出(i.el-icon-download)
# → 正在导出中(.el-loading-mask) → 成功提示(.el-message-box 导出任务建立成功) → 确定
task_success = False
for major_attempt in range(3):
print(f" >> 正在打开数据导出面板 (尝试 {major_attempt + 1}/3)...")
@@ -285,74 +406,61 @@ def yunda_expected_download_impl(page):
state="visible", timeout=15000
)
export_frame = ws_frame.frame_locator('iframe[name="target1"]')
export_frame = _resolve_export_frame(ws_frame)
try:
# 校验字段列表是否加载完成(以“交接单号”为标志)
export_frame.get_by_text("交接单号").first.wait_for(
state="visible", timeout=3000
)
# 校验 Element UI 字段选择区是否加载完成(以「全选」按钮就绪为标志)
export_frame.locator(
"button.el-button", has_text="全选"
).first.wait_for(state="visible", timeout=8000)
except Exception:
print(" ⚠️ 字段列表未加载,关闭面板后重试...")
print(" ⚠️ 导出面板字段区未加载,关闭面板后重试...")
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(1000)
continue
export_frame.locator(".allRight").click()
print(" -> 全选字段并向右转移...")
export_frame.locator("button.el-button", has_text="全选").first.click()
page.wait_for_timeout(400)
export_frame.locator(
"button.el-button:has(i.el-icon-d-arrow-right)"
).first.click()
page.wait_for_timeout(500)
inner_success = False
needs_reopen = False
print(" -> 正在提交导出任务...")
for attempt in range(4):
export_frame.locator("#submitbutton", has_text="导出数据").click()
confirm_link = export_frame.get_by_role("link", name="确定")
try:
confirm_link.wait_for(state="visible", timeout=6000)
if export_frame.get_by_text("导出任务建立成功").is_visible():
print(" ✅ 导出任务已建立成功。")
confirm_link.click()
inner_success = True
break
elif export_frame.get_by_text(
"请选择格式相应的导出字段"
).is_visible():
confirm_link.click()
page.wait_for_timeout(500)
export_frame.locator(
"button.el-button:has(i.el-icon-download)"
).first.click()
# 区分:字段漏选 还是 字段列表消失
if export_frame.get_by_text("交接单号").first.is_visible():
print(
" ⚠️ 检测到未选择字段(字段列表仍在),重新点击全选..."
)
export_frame.locator(".allRight").click()
page.wait_for_timeout(500)
else:
print(
" ⚠️ 字段列表异常消失,重新打开导出面板..."
)
needs_reopen = True
break # 跳出内层循环,重新打开面板
else:
confirm_link.click()
page.wait_for_timeout(1000)
# 等待「正在导出中」遮罩出现并消失
loading = export_frame.locator(".el-loading-mask").first
try:
loading.wait_for(state="visible", timeout=5000)
loading.wait_for(state="hidden", timeout=60000)
except Exception:
page.wait_for_timeout(1000)
pass
# 等待结果提示并判定
inner_success = False
try:
msg_box = export_frame.locator(".el-message-box.my-alert").first
msg_box.wait_for(state="visible", timeout=30000)
if msg_box.get_by_text("导出任务建立成功").is_visible():
print(" ✅ 导出任务已建立成功。")
inner_success = True
else:
print(" ⚠️ 导出结果提示非成功状态,将重试。")
msg_box.locator("button.el-button", has_text="确定").first.click()
page.wait_for_timeout(500)
except Exception:
print(" ⚠️ 未检测到导出结果提示,将重试。")
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(500)
if inner_success:
task_success = True
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(500)
break # 跳出外层循环,继续后续步骤
elif needs_reopen:
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(1000)
continue # 重新打开面板
else:
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(1000)
continue
break
if not task_success:
raise RuntimeError("多次重试后仍未能建立应到数据离线任务。")
@@ -385,18 +493,19 @@ def yunda_expected_download_impl(page):
return False
def yunda_actual_download(page):
def yunda_actual_download(page, force=False, date=None):
"""韵达:实到货物数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry(
"韵达",
"实到",
lambda: yunda_actual_download_impl(page),
lambda: yunda_actual_download_impl(page, date=date),
lambda: yunda_reset(page),
page=page,
)
def yunda_actual_download_impl(page):
def yunda_actual_download_impl(page, date=None):
"""韵达:实到货物数据下载(单次执行,无重试;供自动化测试探测原始结果用)。"""
print("\n▶ 开始执行【韵达 - 实到货物数据下载】任务...")
@@ -427,30 +536,33 @@ def yunda_actual_download_impl(page):
offset = state_store.get_offset("韵达", "actual")
today = datetime.now()
if date:
target = datetime.strptime(date, "%Y-%m-%d")
else:
target = today - timedelta(days=offset)
start_date = target # 单日范围起止同日
today = target # 让下方"截止时间"选择器也指向 target
# 旧版 laydate 的日期格子 td[y][m][d] 用非补零整数值;单日范围起止同日
target_ymd = f"{target.year}-{target.month}-{target.day}"
src = f"指定 {date}" if date else f"偏移 {offset}0=今天"
print(
f">> 设置实到查询日期: [{target.year}-{target.month}-{target.day}]"
f"(偏移 {offset}0=今天)"
f">> 设置实到查询日期: [{target.year}-{target.month}-{target.day}]{src}"
)
print(" >> 正在设定起始时间...")
ws_frame.locator("#startDate").click()
page.wait_for_timeout(400)
box1 = ws_frame.locator("#laydate_box:visible").first
box1.locator(
f"td[y='{start_date.year}'][m='{start_date.month}'][d='{start_date.day}']"
).click()
ws_frame.locator("#laydate_box:visible").first.wait_for(
state="visible", timeout=5000
)
_yunda_pick_laydate_old(ws_frame, page, target_ymd)
page.wait_for_timeout(400)
print(" >> 正在设定截止时间...")
ws_frame.locator("#endDate").click()
page.wait_for_timeout(400)
box2 = ws_frame.locator("#laydate_box:visible").first
box2.locator(
f"td[y='{today.year}'][m='{today.month}'][d='{today.day}']"
).click()
ws_frame.locator("#laydate_box:visible").first.wait_for(
state="visible", timeout=5000
)
_yunda_pick_laydate_old(ws_frame, page, target_ymd)
page.wait_for_timeout(500)
print(" >> 正在变更扫描类型为【到件】...")
@@ -497,8 +609,10 @@ def yunda_actual_download_impl(page):
).click()
return
# 导出弹窗双层重试:外层重新打开面板,内层重新提交
# 区分字段漏选(补点全选)与字段列表消失(重新打开面板)。
# 导出弹窗重试:外层重新打开面板(最多 3 次)
# 韵达站点已将导出面板从 jQuery(.allRight/#submitbutton) 改版为 Element UI
# 全选(button“全选”) → 向右转移(i.el-icon-d-arrow-right) → 导出(i.el-icon-download)
# → 正在导出中(.el-loading-mask) → 成功提示(.el-message-box 导出任务建立成功) → 确定
print(">> 正在发起导出...")
task_success = False
@@ -512,69 +626,58 @@ def yunda_actual_download_impl(page):
export_frame = ws_frame.frame_locator('iframe[name="myFrame"]')
try:
# 校验字段列表是否加载完成(以“扫描类型”为标志)
export_frame.get_by_text("扫描类型").first.wait_for(
state="visible", timeout=3000
)
# 校验 Element UI 字段选择区是否加载完成(以「全选」按钮就绪为标志)
export_frame.locator(
"button.el-button", has_text="全选"
).first.wait_for(state="visible", timeout=8000)
except Exception:
print(" ⚠️ 字段列表未加载,关闭面板后重试...")
print(" ⚠️ 导出面板字段区未加载,关闭面板后重试...")
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(1000)
continue
export_frame.locator(".allRight").click()
print(" -> 全选字段并向右转移...")
export_frame.locator("button.el-button", has_text="全选").first.click()
page.wait_for_timeout(400)
export_frame.locator(
"button.el-button:has(i.el-icon-d-arrow-right)"
).first.click()
page.wait_for_timeout(500)
inner_success = False
needs_reopen = False
print(" -> 正在提交导出任务...")
for attempt in range(4):
export_frame.locator("#submitbutton", has_text="导出数据").click()
confirm_link = export_frame.get_by_role("link", name="确定")
try:
confirm_link.wait_for(state="visible", timeout=6000)
if export_frame.get_by_text("导出任务建立成功").is_visible():
print(" ✅ 导出任务已建立成功。")
confirm_link.click()
inner_success = True
break
elif export_frame.get_by_text(
"请选择格式相应的导出字段"
).is_visible():
confirm_link.click()
page.wait_for_timeout(500)
export_frame.locator(
"button.el-button:has(i.el-icon-download)"
).first.click()
# 区分:字段漏选 还是 字段列表消失
if export_frame.get_by_text("扫描类型").first.is_visible():
print(
" ⚠️ 检测到未选择字段(字段列表仍在),重新点击全选..."
)
export_frame.locator(".allRight").click()
page.wait_for_timeout(500)
else:
print(" ⚠️ 字段列表异常消失,重新打开导出面板...")
needs_reopen = True
break
else:
confirm_link.click()
page.wait_for_timeout(1000)
# 等待「正在导出中」遮罩出现并消失
loading = export_frame.locator(".el-loading-mask").first
try:
loading.wait_for(state="visible", timeout=5000)
loading.wait_for(state="hidden", timeout=60000)
except Exception:
page.wait_for_timeout(1000)
pass
# 等待结果提示并判定
inner_success = False
try:
msg_box = export_frame.locator(".el-message-box.my-alert").first
msg_box.wait_for(state="visible", timeout=30000)
if msg_box.get_by_text("导出任务建立成功").is_visible():
print(" ✅ 导出任务已建立成功。")
inner_success = True
else:
print(" ⚠️ 导出结果提示非成功状态,将重试。")
msg_box.locator("button.el-button", has_text="确定").first.click()
page.wait_for_timeout(500)
except Exception:
print(" ⚠️ 未检测到导出结果提示,将重试。")
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(500)
if inner_success:
task_success = True
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(500)
break
elif needs_reopen:
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(1000)
continue
else:
ws_frame.locator(".layui-layer-close1").click()
page.wait_for_timeout(1000)
continue
if not task_success:
raise RuntimeError("多次重试后仍未能建立实到数据离线任务。")

View File

@@ -1,20 +1,21 @@
# site_zto.py
# sites/zto.py
import os
import re
import time
import yaml
from datetime import datetime
from datetime import datetime, timedelta
import pandas as pd
from paths import DOWNLOAD_DIR, CONFIG_PATH
import state_store
from inbound_verify.paths import DOWNLOAD_DIR, CONFIG_PATH
from inbound_verify import state_store
def with_retry(site_name, label, flow, reset, max_attempts=3):
def with_retry(site_name, label, flow, reset, max_attempts=3, page=None):
"""异常兜底flow 失败 → 重置回初始态 → 重试,最多 max_attempts 次(含首次)。
每次失败都重置含最终放弃那一次既是重试前的清场也保证最终放弃时环境干净
最后一次失败时重置前自动截图保存到 logs/screenshots/供问题排查
flow 为零参可调用返回 False 视为失败其余视为成功
返回 True=最终成功False=重试耗尽放弃供调度层判断任务成败
"""
@@ -28,6 +29,13 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return True
except Exception as e:
print(f"⚠️ 【{site_name}-{label}】第 {attempt}/{max_attempts} 次失败: {e}")
if attempt == max_attempts and page is not None:
try:
from inbound_verify.runtime import capture_error_screenshot
capture_error_screenshot(page, site_name, label, attempt, str(e))
except Exception:
pass
print(f" → 重置【{site_name}】到初始态,清理环境 ...")
try:
reset()
@@ -41,7 +49,7 @@ def with_retry(site_name, label, flow, reset, max_attempts=3):
return False
# 站点首页 URL异常兜底重置用也供 main_router 的 SITES_CONFIG 引用(单一来源)
# 站点首页 URL异常兜底重置用也供 runtime 的 SITES_CONFIG 引用(单一来源)
HOME_URL = "https://ws.zto56.com/"
@@ -93,6 +101,52 @@ def _dom_click(locator):
)
def _zto_compute_target_time(offset):
"""直接从 Python datetime 计算目标日期的毫秒级时间戳(本地时区零点)。
不再从 DOM real-today 元素读取 time 属性避免双月视图下 real-today
同时出现在 month1隐藏 ghost cell month2可见导致 .first 取到隐藏元素
"""
target_date = datetime.now().date() - timedelta(days=offset)
target_dt = datetime(target_date.year, target_date.month, target_date.day)
return int(target_dt.timestamp() * 1000)
def _zto_find_visible_day(frame_locator, target_time):
"""在双月日期控件中查找可见的日期格子。
jQuery Date Range Picker 双月视图下同一天可能出现在两个面板中
- month1左面板的溢出 ghost celldisplay:none不可见
- month2右面板的正常 cell可见
同一日期在 DOM 中可能有毫秒级差异零点 vs 23:59:59遍历匹配并返回
第一个 visible 无可见匹配返回 None
"""
# 尝试两个时间变体:零点 和 23:59:59部分 checked/selected 格用后者)
for time_variant in (target_time, target_time + 86399000):
sel = f"td div.day[time='{time_variant}']"
cells = frame_locator.locator(sel)
count = cells.count()
for i in range(count):
if cells.nth(i).is_visible():
return cells.nth(i)
return None
def _zto_flip_to_target_month(frame_locator, page, target_time, max_flips=12):
"""中通日历(jQuery-Date-Range-Picker 双月视图)跨月导航:目标日期不在当前视窗时,
循环点 .prev 把目标月翻进视窗 _zto_find_visible_day 判可见跳过隐藏 ghost cell
返回 True 若目标格子最终可见"""
for _ in range(max_flips):
if _zto_find_visible_day(frame_locator, target_time) is not None:
return True
frame_locator.locator(".date-picker-wrapper .prev").first.evaluate(
"el => el.click()"
)
page.wait_for_timeout(450)
return _zto_find_visible_day(frame_locator, target_time) is not None
def zto_smart_menu_click(page, menu_path):
"""中通菜单导航"""
print(f">> 正在导航: {' -> '.join(menu_path)}")
@@ -109,18 +163,19 @@ def zto_smart_menu_click(page, menu_path):
page.wait_for_timeout(1000)
def zto_expected_download(page):
def zto_expected_download(page, force=False, date=None):
"""中通:应到货物数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry(
"中通",
"应到",
lambda: zto_expected_download_impl(page),
lambda: zto_expected_download_impl(page, force=force, date=date),
lambda: zto_reset(page),
page=page,
)
def zto_expected_download_impl(page):
def zto_expected_download_impl(page, force=False, date=None):
"""中通:应到货物数据下载(单次执行,无重试;供自动化测试探测原始结果用)。"""
print("\n▶ 开始执行【中通 - 应到货物数据下载】任务...")
@@ -144,33 +199,35 @@ def zto_expected_download_impl(page):
# 读取服务端日期偏移0=今天1=昨天…),单日:起止同日
offset = state_store.get_offset("中通")
if date:
# 指定日期:折算成相对今天的有效偏移,复用下方 target_time 计算与跨月翻月
target_date = datetime.strptime(date, "%Y-%m-%d").date()
offset = (datetime.now().date() - target_date).days
print(f">> 正在设定查询日期: 指定日期 {date}(折算偏移 {offset}...")
else:
print(f">> 正在设定查询日期: 偏移 {offset}0=今天)...")
ewb_frame.locator("#beginDate").click()
page.wait_for_timeout(500)
today_cell = ewb_frame.locator("td div.day.real-today").first
today_cell.wait_for(state="visible")
# 直接从 Python datetime 计算目标时间戳,不再依赖 DOM real-today双月视图
# 下 real-today 可能同时出现在 month1 隐藏 ghost cell 和 month2 可见 cell
# .first 会取到隐藏的那个导致 wait_for(visible) 超时)。
target_time = _zto_compute_target_time(offset)
target_cell = _zto_find_visible_day(ewb_frame, target_time)
today_time_str = today_cell.get_attribute("time")
if today_time_str:
today_time = int(today_time_str)
target_time = today_time - offset * 86400000
target_cell = ewb_frame.locator(f"td div.day[time='{target_time}']").first
if target_cell is None:
print(" 目标日期不在当前视窗,正在翻月导航 ...")
if not _zto_flip_to_target_month(ewb_frame, page, target_time):
raise RuntimeError(f"翻月后仍无法定位目标日期格子(time={target_time})")
target_cell = _zto_find_visible_day(ewb_frame, target_time)
if target_cell is None:
raise RuntimeError(f"翻月后仍无法定位目标日期格子(time={target_time})")
if target_cell.is_visible():
target_cell.click()
# 日期格子用 _dom_click 直接派发事件:.click() 会先 hover 格子,触发
# "范围长度"提示气泡(.date-range-length-tip)盖住格子导致点击被遮挡超时。
_dom_click(target_cell)
page.wait_for_timeout(300)
target_cell.click()
else:
print(" ⚠️ 偏移日期不在当前日历视窗内,自动降级为查询当天。")
today_cell.click()
page.wait_for_timeout(300)
today_cell.click()
else:
today_cell.click()
page.wait_for_timeout(300)
today_cell.click()
_dom_click(target_cell)
page.wait_for_timeout(500)
@@ -240,6 +297,19 @@ def zto_expected_download_impl(page):
count = main_rows.count()
print(f">> 共发现 {count} 个交接单需要导出。")
# 【去重】加载本站已落库交接单号force=True 或查询失败时 existing=空集(不去重)
if force:
existing = set()
print(">> [去重] 强制重下,跳过去重。")
else:
try:
from inbound_verify import store
existing = store.get_existing_handover_nos("中通")
except Exception as _e:
existing = set()
print(f">> [去重] 加载已落库交接单号失败,本次不去重: {_e}")
for i in range(count):
print(f" ⏳ 正在处理第 {i+1}/{count} 个交接单...")
row = ewb_frame.locator(
@@ -251,6 +321,11 @@ def zto_expected_download_impl(page):
handover_no = match.group(0) if match else raw_text.strip()
print(f" -> 当前交接单号:{handover_no}")
# 【去重】已落库的交接单号不再提交导出任务(不双击、不 append export_times
if handover_no in existing:
print(f" ⏭️ 交接单号 {handover_no} 已落库,跳过提交导出。")
continue
row.dblclick()
ewb_frame.locator("#datagrid2").get_by_text("运单号").wait_for(
@@ -316,6 +391,11 @@ def zto_expected_download_impl(page):
except Exception as e:
print(f" ⚠️ 关闭【进站交接单查询】标签页时出错: {e}")
# 【去重兜底】全部已落库/无数据 → 无导出任务,标签页已关,直接结束不进轮询
if not export_times:
print(">> 本次无新交接单需导出(全部已落库或无数据),结束。")
return True
# 交由统一的轮询下载流程处理
_zto_poll_and_download_tasks(
page,
@@ -330,18 +410,19 @@ def zto_expected_download_impl(page):
return False
def zto_actual_download(page):
def zto_actual_download(page, force=False, date=None):
"""中通:实到货物数据下载(内部含异常兜底重试,路由层无感)。"""
return with_retry(
"中通",
"实到",
lambda: zto_actual_download_impl(page),
lambda: zto_actual_download_impl(page, date=date),
lambda: zto_reset(page),
page=page,
)
def zto_actual_download_impl(page):
def zto_actual_download_impl(page, date=None):
"""中通:实到货物数据下载(单次执行,无重试;供自动化测试探测原始结果用)。"""
print("\n▶ 开始执行【中通 - 实到货物数据下载】任务...")
@@ -365,34 +446,34 @@ def zto_actual_download_impl(page):
# 2. 读取服务端日期偏移0=今天1=昨天…),单日:起止同日
offset = state_store.get_offset("中通", "actual")
if date:
target_date = datetime.strptime(date, "%Y-%m-%d").date()
offset = (datetime.now().date() - target_date).days
print(f">> 正在设定查询日期: 指定日期 {date}(折算偏移 {offset}...")
else:
print(f">> 正在设定查询日期: 偏移 {offset}0=今天)...")
arr_frame.locator("#daterange").click()
page.wait_for_timeout(500)
today_cell = arr_frame.locator("td div.day.real-today").first
today_cell.wait_for(state="visible")
# 直接从 Python datetime 计算目标时间戳,不再依赖 DOM real-today双月视图
# 下 real-today 可能同时出现在 month1 隐藏 ghost cell 和 month2 可见 cell
# .first 会取到隐藏的那个导致 wait_for(visible) 超时)。
target_time = _zto_compute_target_time(offset)
target_cell = _zto_find_visible_day(arr_frame, target_time)
today_time_str = today_cell.get_attribute("time")
if today_time_str:
today_time = int(today_time_str)
target_time = today_time - offset * 86400000
target_cell = arr_frame.locator(f"td div.day[time='{target_time}']").first
if target_cell is None:
print(" 目标日期不在当前视窗,正在翻月导航 ...")
if not _zto_flip_to_target_month(arr_frame, page, target_time):
raise RuntimeError(f"翻月后仍无法定位目标日期格子(time={target_time})")
target_cell = _zto_find_visible_day(arr_frame, target_time)
if target_cell is None:
raise RuntimeError(f"翻月后仍无法定位目标日期格子(time={target_time})")
# 日期格子用 _dom_click 直接派发事件Playwright 的 .click() 会先 hover 格子,
# 触发范围长度提示气泡(.date-range-length-tip)盖住格子,导致点击被判遮挡而超时。
if target_cell.is_visible():
# 触发"范围长度"提示气泡(.date-range-length-tip)盖住格子,导致点击被判遮挡而超时。
_dom_click(target_cell)
page.wait_for_timeout(300)
_dom_click(target_cell)
else:
_dom_click(today_cell)
page.wait_for_timeout(300)
_dom_click(today_cell)
else:
_dom_click(today_cell)
page.wait_for_timeout(300)
_dom_click(today_cell)
page.wait_for_timeout(500)
@@ -641,7 +722,7 @@ def _zto_poll_and_download_tasks(page, export_times, download_dir, final_filenam
)
# ====================================================================
# 所有目标文件下载完成后,关闭导出任务管理标签页
# 所有目标文件下载完成后,关闭"导出任务管理"标签页
# ====================================================================
print(">> 【导出任务管理】下载完成,正在关闭标签页...")
try:

View File

@@ -7,9 +7,9 @@
import os
import sqlite3
from datetime import datetime
from datetime import datetime, timedelta
from paths import STATE_DB_PATH
from inbound_verify.paths import STATE_DB_PATH
# 登录态枚举
LOGIN_UNKNOWN = "unknown" # 尚未探测过
@@ -32,7 +32,7 @@ def _now():
def init_db():
"""建库建表(幂等)。确保 state 目录存在。"""
os.makedirs(os.path.dirname(STATE_DB_PATH), exist_ok=True)
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
conn.execute("""
CREATE TABLE IF NOT EXISTS site_status (
site TEXT PRIMARY KEY,
@@ -70,9 +70,22 @@ def init_db():
status TEXT,
started_at TEXT,
finished_at TEXT,
error TEXT
error TEXT,
trigger TEXT NOT NULL DEFAULT '',
target_date TEXT NOT NULL DEFAULT '',
force INTEGER NOT NULL DEFAULT 0
)
""")
# 旧库迁移:补触发方式/目标日期/强制重下三列(新库已含;重复添加抛 OperationalError忽略
for _col, _typedef in [
("trigger", "TEXT NOT NULL DEFAULT ''"),
("target_date", "TEXT NOT NULL DEFAULT ''"),
("force", "INTEGER NOT NULL DEFAULT 0"),
]:
try:
conn.execute(f"ALTER TABLE task_history ADD COLUMN {_col} {_typedef}")
except sqlite3.OperationalError:
pass
conn.execute("""
CREATE TABLE IF NOT EXISTS site_config (
site TEXT PRIMARY KEY,
@@ -110,6 +123,32 @@ def init_db():
PRIMARY KEY (site, key)
)
""")
conn.execute("""
CREATE TABLE IF NOT EXISTS ingest_state (
site TEXT,
kind TEXT,
ok INTEGER,
ingested_at TEXT,
count INTEGER,
error TEXT,
PRIMARY KEY (site, kind)
)
""")
# 周期性抓取调度(每 site×kind 一行):取代旧 site_config.schedule_* 的每日单时点。
# enabled=总开关active_start/end=激活时段"HH:MM"(空=不限时段,避免半夜空跑);
# interval_minutes=激活时段内的抓取间隔。百世 kind 固定为 undelivered无应到/实到二分)。
conn.execute("""
CREATE TABLE IF NOT EXISTS fetch_schedule (
site TEXT NOT NULL,
kind TEXT NOT NULL,
enabled INTEGER NOT NULL DEFAULT 0,
active_start TEXT NOT NULL DEFAULT '',
active_end TEXT NOT NULL DEFAULT '',
interval_minutes INTEGER NOT NULL DEFAULT 30,
updated_at TEXT,
PRIMARY KEY (site, kind)
)
""")
conn.commit()
@@ -191,14 +230,14 @@ def _upsert(conn, site, **fields):
def set_login_state(site, logged_in):
"""更新单站登录态。logged_in: bool。"""
state = LOGIN_IN if logged_in else LOGIN_OUT
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
_upsert(conn, site, login_state=state, login_checked_at=_now())
def reset_login_states(sites):
"""启动时把给定站点的登录态重置为 unknown避免显示上一会话的陈旧登录态
登录态是会话级的数据态文件就绪会话无关保留不动心跳就绪后会重新探测"""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
for site in sites:
_upsert(conn, site, login_state=LOGIN_UNKNOWN, login_checked_at="")
@@ -212,15 +251,34 @@ def set_data_state(site, kind, ready, generated_at, business_date=None):
}
if business_date is not None:
fields[f"{kind}_business_date"] = business_date or ""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
_upsert(conn, site, **fields)
def set_business_date(site, kind, business_date):
"""仅写业务日期快照(不碰 ready/generated_at
下载成功钩子用ready 语义已移交入库成功 reset_data_ready / _persist_to_db
下载阶段只记业务日期供前端状态盘显示是哪天的数据
"""
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
_upsert(conn, site, **{f"{kind}_business_date": business_date or ""})
def set_ready(site, kind, ready):
"""仅写就绪态(不碰 business_date/generated_at
供心跳从 ingest_state 派生 ready ready 现为 DB 入库真相的派生视图
非启动重置不读 Excel"""
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
_upsert(conn, site, **{f"{kind}_ready": 1 if ready else 0})
def get_all_status():
"""返回 {site: {各字段}};库不存在则返回 {}"""
if not os.path.exists(STATE_DB_PATH):
return {}
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
rows = conn.execute(
"SELECT site, login_state, login_checked_at, expected_ready, "
"expected_generated_at, expected_business_date, actual_ready, "
@@ -247,6 +305,39 @@ def get_all_status():
}
def set_ingest_state(site, kind, ok, count=0, error=None):
"""记录一次入库结果UPSERT。ok: boolcount: 入库条数error: 失败原因或 None。"""
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
conn.execute(
"INSERT INTO ingest_state (site, kind, ok, ingested_at, count, error) "
"VALUES (?, ?, ?, ?, ?, ?) "
"ON CONFLICT(site, kind) DO UPDATE SET "
"ok=excluded.ok, ingested_at=excluded.ingested_at, "
"count=excluded.count, error=excluded.error",
(site, kind, 1 if ok else 0, _now(), int(count or 0), error or ""),
)
conn.commit()
def get_all_ingest_state():
"""返回 {site: {kind: {ok, ingested_at, count, error}}};库不存在返回 {}"""
if not os.path.exists(STATE_DB_PATH):
return {}
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
rows = conn.execute(
"SELECT site, kind, ok, ingested_at, count, error FROM ingest_state"
).fetchall()
out = {}
for site, kind, ok, ingested_at, count, error in rows:
out.setdefault(site, {})[kind] = {
"ok": bool(ok),
"ingested_at": ingested_at or "",
"count": int(count or 0),
"error": error or "",
}
return out
# ============================ 下载日期偏移site_config============================
MAX_DATE_OFFSET = 30 # 0=今天,最大回溯 30 天
@@ -257,7 +348,7 @@ def get_offset(site, kind="expected"):
col = "expected_offset" if kind == "expected" else "actual_offset"
if not os.path.exists(STATE_DB_PATH):
return 0
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
row = conn.execute(
f"SELECT {col} FROM site_config WHERE site=?", (site,)
).fetchone()
@@ -268,7 +359,7 @@ def set_offset(site, kind, offset):
"""设置单站下载日期偏移kind: 'expected'/'actual'),钳制到 [0, MAX_DATE_OFFSET]。"""
col = "expected_offset" if kind == "expected" else "actual_offset"
offset = max(0, min(MAX_DATE_OFFSET, int(offset)))
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
conn.execute(
f"INSERT INTO site_config (site, {col}, updated_at) VALUES (?, ?, ?) "
f"ON CONFLICT(site) DO UPDATE SET {col}=excluded.{col}, "
@@ -279,10 +370,32 @@ def set_offset(site, kind, offset):
return offset
def resolve_target_date(site, kind, date=None):
"""计算一条任务的目标下载日期YYYY-MM-DD供任务日志展示 / 重试回放)。
date date否则按站点偏移推算 runtime._record_business_date 同源
expected 应到偏移actual 实到偏移百世 undelivered 当天
4 undelivered 跟随应到偏移__compare__ 无数据概念返回 ''"""
if site == "__compare__":
return ""
if date:
return date
today = datetime.now().date()
if kind == "expected":
return (today - timedelta(days=get_offset(site, "expected"))).strftime(
"%Y-%m-%d"
)
if kind == "actual":
return (today - timedelta(days=get_offset(site, "actual"))).strftime("%Y-%m-%d")
if site == "百世":
return today.strftime("%Y-%m-%d")
return (today - timedelta(days=get_offset(site, "expected"))).strftime("%Y-%m-%d")
def set_schedule(site, enabled, time_str):
"""设置单站每日定时下载enabled: booltime_str: 'HH:MM''')。"""
"""【DEPRECATED】旧"每日单时点定时"——已被 fetch_schedule 的周期+激活时段模式取代。
保留死代码以防外部残留调用新代码请用 set_fetch_schedule"""
enabled_int = 1 if enabled else 0
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
conn.execute(
"INSERT INTO site_config (site, schedule_enabled, schedule_time, updated_at) "
"VALUES (?, ?, ?, ?) "
@@ -295,24 +408,129 @@ def set_schedule(site, enabled, time_str):
return bool(enabled_int), (time_str or "")
def get_all_config():
"""返回 {site: {expected_offset, actual_offset, schedule_enabled, schedule_time}}。"""
if not os.path.exists(STATE_DB_PATH):
return {}
with sqlite3.connect(STATE_DB_PATH) as conn:
rows = conn.execute(
"SELECT site, expected_offset, actual_offset, schedule_enabled, schedule_time "
"FROM site_config"
).fetchall()
# ============================ 周期性抓取调度fetch_schedule============================
# 取代旧 site_config.schedule_* 的"每日单时点":每 site×kind 一行,
# 在激活时段 [active_start, active_end) 内按 interval_minutes 周期抓取。
# 应到/实到各自独立配置;百世只有 undelivered站点直供未到明细无应到/实到二分)。
# 各站允许的周期抓取 kind单一来源server.py 复用)
SITE_FETCH_KINDS = {
"顺心": ("expected", "actual"),
"中通": ("expected", "actual"),
"韵达": ("expected", "actual"),
"安能": ("expected", "actual"),
"百世": ("undelivered",),
}
DEFAULT_FETCH_SCHEDULE = {
"enabled": False,
"active_start": "",
"active_end": "",
"interval_minutes": 30,
}
def allowed_kinds(site):
"""该站允许的周期抓取 kind 元组;未知站点返回空元组。"""
return SITE_FETCH_KINDS.get(site, ())
def _fetch_spec(row):
"""把 fetch_schedule 行转成 spec dict。"""
return {
r[0]: {
"expected_offset": int(r[1]),
"actual_offset": int(r[2]),
"schedule_enabled": bool(r[3]),
"schedule_time": r[4] or "",
"enabled": bool(row[0]),
"active_start": row[1] or "",
"active_end": row[2] or "",
"interval_minutes": int(row[3]),
}
def get_fetch_schedule(site, kind):
"""读单 (site,kind) 周期抓取配置;未配置返回 None。"""
if not os.path.exists(STATE_DB_PATH):
return None
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
row = conn.execute(
"SELECT enabled, active_start, active_end, interval_minutes "
"FROM fetch_schedule WHERE site=? AND kind=?",
(site, kind),
).fetchone()
return _fetch_spec(row) if row else None
def set_fetch_schedule(site, kind, enabled, active_start, active_end, interval_minutes):
"""UPSERT 单 (site,kind) 周期抓取配置;返回写入后的 spec dict。"""
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
conn.execute(
"INSERT INTO fetch_schedule "
"(site, kind, enabled, active_start, active_end, interval_minutes, updated_at) "
"VALUES (?, ?, ?, ?, ?, ?, ?) "
"ON CONFLICT(site, kind) DO UPDATE SET "
"enabled=excluded.enabled, active_start=excluded.active_start, "
"active_end=excluded.active_end, interval_minutes=excluded.interval_minutes, "
"updated_at=excluded.updated_at",
(
site,
kind,
1 if enabled else 0,
(active_start or ""),
(active_end or ""),
int(interval_minutes),
_now(),
),
)
conn.commit()
return {
"enabled": bool(enabled),
"active_start": active_start or "",
"active_end": active_end or "",
"interval_minutes": int(interval_minutes),
}
def get_all_fetch_schedules():
"""返回 {site: {kind: spec}};对每个站点的每个 allowed kind 都补齐(缺失用默认值)。
保证前端永远拿到完整 kind 不必前端补默认 spec"""
out = {}
if not os.path.exists(STATE_DB_PATH):
for site, kinds in SITE_FETCH_KINDS.items():
out[site] = {k: dict(DEFAULT_FETCH_SCHEDULE) for k in kinds}
return out
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
rows = conn.execute(
"SELECT site, kind, enabled, active_start, active_end, interval_minutes "
"FROM fetch_schedule"
).fetchall()
by_key = {(r[0], r[1]): _fetch_spec(r[2:]) for r in rows}
for site, kinds in SITE_FETCH_KINDS.items():
out[site] = {
k: by_key.get((site, k), dict(DEFAULT_FETCH_SCHEDULE)) for k in kinds
}
return out
def get_all_config():
"""返回 {site: {expected_offset, actual_offset, fetch_schedules}}。
站点集以 fetch_schedule allowed kinds 为准覆盖全业务站点offsets 缺失默认 0"""
schedules = get_all_fetch_schedules()
offsets = {}
if os.path.exists(STATE_DB_PATH):
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
rows = conn.execute(
"SELECT site, expected_offset, actual_offset FROM site_config"
).fetchall()
offsets = {
r[0]: {"expected_offset": int(r[1]), "actual_offset": int(r[2])}
for r in rows
}
return {
site: {
"expected_offset": offsets.get(site, {}).get("expected_offset", 0),
"actual_offset": offsets.get(site, {}).get("actual_offset", 0),
"fetch_schedules": schedules.get(site, {}),
}
for site in schedules
}
# ============================ 站点键值配置site_settings============================
@@ -323,7 +541,7 @@ def get_setting(site, key):
"""读取单站某个配置值;未设置返回 ''"""
if not os.path.exists(STATE_DB_PATH):
return ""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
row = conn.execute(
"SELECT value FROM site_settings WHERE site=? AND key=?", (site, key)
).fetchone()
@@ -332,7 +550,7 @@ def get_setting(site, key):
def set_setting(site, key, value):
"""设置单站某个配置值upsert"""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
conn.execute(
"INSERT INTO site_settings (site, key, value) VALUES (?, ?, ?) "
"ON CONFLICT(site, key) DO UPDATE SET value=excluded.value",
@@ -345,7 +563,7 @@ def get_site_settings(site):
"""返回单站全部配置 {key: value}。"""
if not os.path.exists(STATE_DB_PATH):
return {}
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
rows = conn.execute(
"SELECT key, value FROM site_settings WHERE site=?", (site,)
).fetchall()
@@ -355,14 +573,42 @@ def get_site_settings(site):
# ============================ 任务历史 ============================
def create_task(site, kind):
"""新建一条 pending 任务,返回其 id。"""
def create_task(site, kind, trigger="manual", target_date="", force=False):
"""新建一条 pending 任务(手动触发),返回其 id。trigger='manual'/'auto'
target_date 为该任务的目标下载日期YYYY-MM-DD可为 ''force 是否强制重下"""
now = _now()
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
cur = conn.execute(
"INSERT INTO task_history (site, kind, status, started_at, finished_at, error) "
"VALUES (?, ?, ?, ?, '', '')",
(site, kind, TASK_PENDING, now),
"INSERT INTO task_history "
"(site, kind, status, started_at, finished_at, error, trigger, target_date, force) "
"VALUES (?, ?, ?, ?, '', '', ?, ?, ?)",
(site, kind, TASK_PENDING, now, trigger, target_date, 1 if force else 0),
)
conn.commit()
return cur.lastrowid
def create_task_if_idle(site, kind, trigger="auto", target_date=""):
"""周期调度专用:若该 (site,kind) 已有 pending/running 任务则返回 None跳过本次周期
否则建一条 pending 任务返回其 id单连接内 check-then-insert SQLite 写锁把竞态压到忽略不计
create_task 的区别手动触发(POST /tasks) create_task用户点的必建周期 job 用本函数
上一次还没跑完时跳过避免同 (site,kind) 任务堆积手动建的任务会让紧随其后的周期 fire
判到 inflight 而跳过天然互斥trigger='auto'target_date 为目标下载日期YYYY-MM-DD"""
now = _now()
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
row = conn.execute(
"SELECT 1 FROM task_history WHERE site=? AND kind=? "
"AND status IN (?, ?) LIMIT 1",
(site, kind, TASK_PENDING, TASK_RUNNING),
).fetchone()
if row:
return None
cur = conn.execute(
"INSERT INTO task_history "
"(site, kind, status, started_at, finished_at, error, trigger, target_date, force) "
"VALUES (?, ?, ?, ?, '', '', ?, ?, 0)",
(site, kind, TASK_PENDING, now, trigger, target_date),
)
conn.commit()
return cur.lastrowid
@@ -371,7 +617,7 @@ def create_task(site, kind):
def update_task(task_id, status, error=None):
"""更新任务状态。终态(success/no_data/failed)写入 finished_at。"""
finished = _now() if status in (TASK_SUCCESS, TASK_NO_DATA, TASK_FAILED) else ""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
if finished:
conn.execute(
"UPDATE task_history SET status=?, finished_at=?, error=? WHERE id=?",
@@ -387,9 +633,10 @@ def update_task(task_id, status, error=None):
def get_task(task_id):
"""返回单条任务 dict不存在返回 None。"""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
row = conn.execute(
"SELECT id, site, kind, status, started_at, finished_at, error "
"SELECT id, site, kind, status, started_at, finished_at, error, "
"trigger, target_date, force "
"FROM task_history WHERE id=?",
(task_id,),
).fetchone()
@@ -403,14 +650,18 @@ def get_task(task_id):
"started_at": row[4],
"finished_at": row[5],
"error": row[6],
"trigger": row[7],
"target_date": row[8],
"force": bool(row[9]),
}
def list_tasks(limit=20):
"""返回最近 limit 条任务(按 id 倒序)。"""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
rows = conn.execute(
"SELECT id, site, kind, status, started_at, finished_at, error "
"SELECT id, site, kind, status, started_at, finished_at, error, "
"trigger, target_date, force "
"FROM task_history ORDER BY id DESC LIMIT ?",
(limit,),
).fetchall()
@@ -423,6 +674,9 @@ def list_tasks(limit=20):
"started_at": r[4],
"finished_at": r[5],
"error": r[6],
"trigger": r[7],
"target_date": r[8],
"force": bool(r[9]),
}
for r in rows
]
@@ -431,7 +685,7 @@ def list_tasks(limit=20):
def fail_stale_tasks(reason: str = "服务重启,上轮未完成任务,请手动重跑") -> int:
"""worker 启动时调用:把遗留的 pending/running 任务标记为 failed实现重启自愈。
返回被清理的任务数量"""
with sqlite3.connect(STATE_DB_PATH) as conn:
with sqlite3.connect(STATE_DB_PATH, timeout=3.0) as conn:
cur = conn.execute(
"UPDATE task_history SET status=?, finished_at=?, error=? "
"WHERE status IN (?, ?)",

View File

@@ -1,6 +1,6 @@
# -*- coding: utf-8 -*-
"""
db_store.py 到货核销数据持久化PostgreSQL
store.py 到货核销数据持久化PostgreSQL
职责 downloads/ 下各站点下载的应到 / 实到 / 未到 Excel 解析后幂等写入 PostgreSQL
与下载流程解耦本模块只读 downloads/ 现有文件入库不关心谁触发下载下载了几次
@@ -13,10 +13,11 @@ db_store.py — 到货核销数据持久化PostgreSQL
- 单号一律按文本读写dtype=str防长数字被科学计数 / 精度丢失
命令行
python db_store.py createdb 创建数据库幂等
python db_store.py init 建表幂等 CREATE TABLE IF NOT EXISTS
python db_store.py ingest [site] 入库全站或单站幂等 UPSERT
python db_store.py all createdb init 全站 ingest 一条龙
python -m inbound_verify.store createdb 创建数据库幂等
python -m inbound_verify.store init 建表幂等 CREATE TABLE IF NOT EXISTS
python -m inbound_verify.store ingest [site] 入库全站或单站幂等 UPSERT
python -m inbound_verify.store ingest-one <site> <kind> 仅入库指定站/钩子同款路由
python -m inbound_verify.store all createdb init 全站 ingest 一条龙
"""
import os
@@ -28,9 +29,13 @@ import psycopg
import yaml
from psycopg.types.json import Jsonb
from paths import BASE_DIR, CONFIG_PATH, DOWNLOAD_DIR
from inbound_verify.paths import BASE_DIR, CONFIG_PATH, DOWNLOAD_DIR
from inbound_verify.domain import (
BAISHI_FILE,
_site_cfg,
) # 站点 / 文件名配置(单一来源)
import expected_undelivered as eu # 复用站点 / 文件名 / 列映射 / 基号口径(单一来源
from inbound_verify import compare # _read_business_dates比对侧业务日期读取
SCHEMA_PATH = os.path.join(BASE_DIR, "schema.sql")
@@ -57,12 +62,15 @@ def _load_pg_config():
"password": pg.get("password", ""),
"dbname": pg.get("dbname", "CQHXDB"),
"schema": pg.get("schema", "inbound_verify"),
"auto_ingest": bool(pg.get("auto_ingest", True)),
"connect_timeout_seconds": int(pg.get("connect_timeout_seconds", 5)),
}
def _connect(dbname):
"""用关键字参数连接(避开 conninfo 对密码特殊字符的解析)。
options search_path 到专用 schema使无 schema 限定的表名解析到该 schema"""
options search_path 到专用 schema + 会话级 statement_timeout=30s
cpolar 隧道上防失控查询connect_timeout 守连接阶段"""
c = _load_pg_config()
return psycopg.connect(
host=c["host"],
@@ -70,10 +78,17 @@ def _connect(dbname):
dbname=dbname,
user=c["user"],
password=c["password"],
options=f"-c search_path={c['schema']}",
options=f"-c search_path={c['schema']} -c statement_timeout=30s",
connect_timeout=c["connect_timeout_seconds"],
)
def ingest_enabled():
"""是否启用下载后自动入库config.yaml postgres.auto_ingest默认 True
runtime 钩子判定开关避免它伸手进 _load_pg_config"""
return _load_pg_config()["auto_ingest"]
# ============================== 建库 / 建表 ==============================
@@ -105,7 +120,7 @@ def init_schema():
# ============================== 解析辅助 ==============================
# 实到单号列 / 基号列映射(口径取自 expected_undelivered
# 实到单号列 / 基号列映射(口径取自 domain与 STATIONS 对齐
# piece = 实到表里「每件」的单号列(扫描单号 / 子单号 / 复合串)
# waybill = 与应到运单号对齐的干净列(中通无干净列,由 piece 复合串 v[:-8] 推导)
# scan_time = 扫描时间列(缺失则不填,原始值仍在 raw
@@ -208,7 +223,7 @@ def _raw_row(row):
def _read_business_dates():
"""从状态库读各站本次业务日期(与报告口径一致;读不到返回空 dict"""
try:
return eu._read_business_dates(ALL_SITES + ["百世"]) or {}
return compare._read_business_dates(ALL_SITES + ["百世"]) or {}
except Exception as e:
print(f">> [warn] 读取业务日期失败(不影响入库): {e}")
return {}
@@ -253,13 +268,25 @@ _SQL_UNDELIVERED = """
ingested_at = now()
"""
_SQL_BAISHI_DAILY_STATS = """
INSERT INTO baishi_daily_stats
(site, business_date, expected_pieces, arrived_pieces, undelivered_pieces, raw)
VALUES (%s,%s,%s,%s,%s,%s)
ON CONFLICT (site, business_date) DO UPDATE SET
expected_pieces = COALESCE(EXCLUDED.expected_pieces, baishi_daily_stats.expected_pieces),
arrived_pieces = COALESCE(EXCLUDED.arrived_pieces, baishi_daily_stats.arrived_pieces),
undelivered_pieces = COALESCE(EXCLUDED.undelivered_pieces, baishi_daily_stats.undelivered_pieces),
raw = EXCLUDED.raw,
ingested_at = now()
"""
# ============================== 入库 ==============================
def _ingest_expected(cur, site, business_date):
"""入库单站应到(运单级,按 waybill_no 去重 keep-first 后 UPSERT"""
cfg = eu._site_cfg(site)
cfg = _site_cfg(site)
path = os.path.join(DOWNLOAD_DIR, cfg["exp"])
if not os.path.exists(path):
print(f" [跳过] {site} 应到:文件不存在 {cfg['exp']}")
@@ -291,7 +318,7 @@ def _ingest_expected(cur, site, business_date):
def _ingest_actual(cur, site):
"""入库单站实到(扫描件级,按 piece_no UPSERT"""
cfg = eu._site_cfg(site)
cfg = _site_cfg(site)
path = os.path.join(DOWNLOAD_DIR, cfg["act"])
if not os.path.exists(path):
print(f" [跳过] {site} 实到:文件不存在 {cfg['act']}")
@@ -299,9 +326,10 @@ def _ingest_actual(cur, site):
cm = ACTUAL_COLMAP[site]
df = pd.read_excel(path, dtype=str).fillna("")
if site == "韵达":
# 韵达业务清洗:抛弃「交接单号」为空的行(派件/签收等其他扫描无交接单号
# 韵达业务清洗:保留「交接单号」为空的行(到/接件扫描
# 抛弃「交接单号」不为空的行(派件/签收等,属重复数据)。
# 再按子单号去重一件多扫只留一条清洗后子单号已天然唯一drop 为保险)。
df = df[df["交接单号"].astype(str).str.strip() != ""]
df = df[df["交接单号"].astype(str).str.strip() == ""]
df = df.drop_duplicates(subset=[cm["piece"]], keep="last")
rows = []
for r in df.to_dict("records"):
@@ -330,9 +358,9 @@ def _ingest_actual(cur, site):
def _ingest_undelivered_baishi(cur):
"""入库百世应到未到明细(子单级,按 (site, piece_no) UPSERT"""
path = os.path.join(DOWNLOAD_DIR, eu.BAISHI_FILE)
path = os.path.join(DOWNLOAD_DIR, BAISHI_FILE)
if not os.path.exists(path):
print(f" [跳过] 百世 未到:文件不存在 {eu.BAISHI_FILE}")
print(f" [跳过] 百世 未到:文件不存在 {BAISHI_FILE}")
return 0
df = pd.read_excel(path, dtype=str).fillna("")
rows = []
@@ -353,6 +381,34 @@ def _ingest_undelivered_baishi(cur):
return len(rows)
def upsert_baishi_daily_stats(exp, arr, business_date=None):
"""直接落库百世当日应到/实到基数(应扫/已扫,站级日聚合)。
baishi 下载时抓到基数后直接调用一步落库不绕 state_storestore
business_date 默认今天百世固定当天best-effort失败只告警不影响下载流程"""
biz = business_date or date.today()
if exp is None and arr is None:
return
undel = (exp - arr) if (exp is not None and arr is not None) else None
try:
with _connect(_load_pg_config()["dbname"]) as conn:
with conn.cursor() as cur:
cur.execute(
_SQL_BAISHI_DAILY_STATS,
(
"百世",
biz,
exp,
arr,
undel,
Jsonb({"expected": exp, "arrived": arr, "undelivered": undel}),
),
)
conn.commit()
print(f" [基数] 百世 {biz}: 应扫 {exp} / 已扫 {arr} / 未扫 {undel}")
except Exception as e:
print(f" [基数] 百世 {biz} 入库失败(不影响下载): {e}")
def ingest(site=None):
"""入库:指定 site 则单站(百世只入未到),否则全站。返回总条数。"""
dates = _read_business_dates()
@@ -373,10 +429,109 @@ def ingest(site=None):
return total
def ingest_task(site, kind):
"""按 (site, kind) 入库本次刚下载的文件(幂等 UPSERT返回总条数。
ingest(site) 的区别只入本次刷新的那一类避免重读写另一类文件同步钩子里减少阻塞
kind 路由
expected/actual 各入其列
undelivered 百世 入未到
undelivered 4 _site_undelivered_handler 内部连带下了 expected+actual故入两者
__compare__ / 其它组合 返回 0
"""
if site == "__compare__":
return 0
# 百世无应到/实到(只有站点直供的未到);非 undelivered 直接返回 0避免
# _ingest_expected/_ingest_actual 走到 _site_cfg(百世)=None 而 TypeError。
if site == "百世" and kind != "undelivered":
return 0
dates = _read_business_dates()
total = 0
with _connect(_load_pg_config()["dbname"]) as conn:
with conn.cursor() as cur:
if kind == "expected":
total += _ingest_expected(cur, site, dates.get(site))
elif kind == "actual":
total += _ingest_actual(cur, site)
elif kind == "undelivered":
if site == "百世":
total += _ingest_undelivered_baishi(cur)
else: # 顺心/中通/韵达/安能
total += _ingest_expected(cur, site, dates.get(site))
total += _ingest_actual(cur, site)
# 其它组合(如 百世/expected正常不经钩子触发防御性返回 0
conn.commit()
return total
def get_existing_handover_nos(site):
"""查该站点已落库的交接单号集合expected_record.handover_no
"提交导出任务前"去重已落库的交接单号不再重复提交导出任务
PG 不可用cpolar 抖动等时返回空集 + 告警调用方按"未确认存在"处理
继续提交导出UPSERT 兜底绝不因去重查询失败而漏数据"""
try:
with _connect(_load_pg_config()["dbname"]) as conn:
with conn.cursor() as cur:
cur.execute(
"SELECT handover_no FROM expected_record "
"WHERE site=%s AND handover_no IS NOT NULL AND handover_no <> ''",
(site,),
)
return {str(r[0]).strip() for r in cur.fetchall()}
except Exception as e:
print(f">> [去重] 查询已落库交接单号失败({site}),本次不去重: {e}")
return set()
# ============================== PG 数据存在性查询 ==============================
def has_data(site, kind, target_date):
"""查询 PG指定站点在 target_date 是否有业务数据。
target_date: str 'YYYY-MM-DD' date 对象
返回 (has_rows: bool, count: int)
PG 不可达时返回 (False, 0)不抛异常调用方按未确认存在处理
kind 路由
expected expected_record (business_date)
actual actual_record (scan_time::date)
undelivered 百世: baishi_daily_stats4 : 不单独查由调用方 expectedactual 派生
"""
if site == "百世" and kind == "undelivered":
sql = (
"SELECT COUNT(*) FROM baishi_daily_stats"
" WHERE site = %s AND business_date = %s"
)
params = (site, target_date)
elif kind == "expected":
sql = (
"SELECT COUNT(*) FROM expected_record"
" WHERE site = %s AND business_date = %s"
)
params = (site, target_date)
elif kind == "actual":
sql = (
"SELECT COUNT(*) FROM actual_record"
" WHERE site = %s AND scan_time::date = %s"
)
params = (site, target_date)
else:
return (False, 0)
try:
with _connect(_load_pg_config()["dbname"]) as conn:
with conn.cursor() as cur:
cur.execute(sql, params)
row = cur.fetchone()
cnt = int(row[0]) if row else 0
return (cnt > 0, cnt)
except Exception as e:
print(f">> [状态] PG 查询 {site}/{kind}/{target_date} 失败: {e}")
return (False, 0)
# ============================== 命令行 ==============================
def _cli():
def main():
cmd = sys.argv[1] if len(sys.argv) > 1 else "all"
site = sys.argv[2] if len(sys.argv) > 2 else None
if cmd == "createdb":
@@ -389,10 +544,19 @@ def _cli():
create_database()
init_schema()
ingest()
elif cmd == "ingest-one":
kind = sys.argv[3] if len(sys.argv) > 3 else None
if not site or kind not in ("expected", "actual", "undelivered"):
print(
"用法: python -m inbound_verify.store ingest-one <site> <expected|actual|undelivered>"
)
sys.exit(1)
total = ingest_task(site, kind)
print(f">> [ingest-one] {site}/{kind} 入库 {total}")
else:
print(__doc__)
sys.exit(1)
if __name__ == "__main__":
_cli()
main()

28
pyproject.toml Normal file
View File

@@ -0,0 +1,28 @@
[build-system]
requires = ["setuptools>=68"]
build-backend = "setuptools.build_meta"
[project]
name = "inbound-verify"
version = "0.1.0"
description = "物流到货数据自动下载与应到未到核对工具"
requires-python = ">=3.10"
dependencies = [
"pandas>=2.0.0",
"playwright>=1.40.0",
"openpyxl>=3.1.0",
"PyYAML>=6.0",
"websocket-client>=1.0.0",
"fastapi>=0.110.0",
"uvicorn>=0.27.0",
"apscheduler>=3.10.0",
"psycopg[binary]>=3.1",
]
[project.scripts]
inbound-verify = "inbound_verify.cli.router:main"
inbound-verify-server = "inbound_verify.cli.server:main"
inbound-verify-db = "inbound_verify.store:main"
[tool.setuptools.packages.find]
include = ["inbound_verify*"]

View File

@@ -24,6 +24,7 @@ CREATE TABLE IF NOT EXISTS expected_record (
UNIQUE (site, waybill_no)
);
CREATE INDEX IF NOT EXISTS idx_expected_site_date ON expected_record (site, business_date);
CREATE INDEX IF NOT EXISTS idx_expected_handover ON expected_record (site, handover_no);
-- 实到货物(扫描件级:一扫描一行;每扫描一件系统生成一个单号)
CREATE TABLE IF NOT EXISTS actual_record (
@@ -52,3 +53,16 @@ CREATE TABLE IF NOT EXISTS undelivered_record (
ingested_at TIMESTAMPTZ NOT NULL DEFAULT now(),
UNIQUE (site, piece_no)
);
-- 百世日聚合(应扫/已扫基数:站级日聚合,区别于运单级/件级/子单级表)
CREATE TABLE IF NOT EXISTS baishi_daily_stats (
id BIGSERIAL PRIMARY KEY,
site TEXT NOT NULL, -- 百世
business_date DATE NOT NULL, -- 业务日期(百世固定当天)
expected_pieces INTEGER, -- 应扫(应到基数)
arrived_pieces INTEGER, -- 已扫(实到基数)
undelivered_pieces INTEGER, -- 未扫(=应扫-已扫,任一缺失则 NULL
raw JSONB NOT NULL, -- 原始抓取值
ingested_at TIMESTAMPTZ NOT NULL DEFAULT now(),
UNIQUE (site, business_date)
);

291
server.py
View File

@@ -1,291 +0,0 @@
# server.py
#
# 服务模式入口FastAPI主线程uvicorn/asyncio+ Playwright worker独立线程
# 客户端经 HTTP 触发任务、查状态、下载数据worker 串行执行任务并跑心跳。
#
# 线程模型(关键):
# - 主线程uvicorn + FastAPI。路由【绝不】访问 Playwright 对象,只经
# task_queue投递任务+ state_store查状态/任务)+ 文件系统(下载数据)。
# - worker 线程runtime.launch_and_preparesync_playwright 在此)+ 任务循环,
# 独占所有 page 操作;与主线程仅经 Queue + SQLite 通信。
# 违反"路由不碰 Playwright"会崩sync 对象跨线程访问)。
#
# 运行python server.py (默认监听 0.0.0.0:8000
import os
import queue
import threading
import time
from contextlib import asynccontextmanager
from typing import Dict, Optional
import uvicorn
from apscheduler.schedulers.background import BackgroundScheduler
from apscheduler.triggers.cron import CronTrigger
from fastapi import FastAPI, HTTPException
from fastapi.responses import FileResponse
from pydantic import BaseModel
from paths import DOWNLOAD_DIR, OUTPUT_DIR
import state_store
from runtime import (
HEARTBEAT_INTERVAL,
TASK_HANDLERS,
dispatch_task,
launch_and_prepare,
run_heartbeat,
)
# 全部站点;百世固定下载当天,不可配置偏移
ALL_SITES = ["顺心", "百世", "中通", "韵达", "安能"]
CONFIGURABLE_SITES = {"顺心", "中通", "韵达", "安能"}
# 任务队列:元素 (task_id, task_spec)。主线程投递worker 消费。
task_queue: "queue.Queue" = queue.Queue()
# worker 运行状态主线程只读worker 写)
worker_state = {
"ctx": None,
"stop": False,
"thread": None,
"ready": False, # launch_and_prepare 完成(各站就绪,可接任务)
"error": None, # worker 启动失败原因
}
# 每日定时下载调度器进程内job 只往 task_queue 投任务,不碰 Playwright
scheduler = BackgroundScheduler(daemon=True)
def _worker_loop():
"""worker 线程:启动 Playwright + 等就绪 + 任务循环(执行任务 + 心跳)。"""
try:
ctx = launch_and_prepare()
worker_state["ctx"] = ctx
worker_state["ready"] = True
# 【P1-2 重启自愈】worker 就绪后清理上轮遗留的 pending/running 僵尸任务
cleaned = state_store.fail_stale_tasks()
if cleaned:
print(f">> [worker] 自愈:清理 {cleaned} 条遗留任务pending/running → failed")
print(">> [worker] 各站就绪,开始接收任务 ...")
except Exception as e:
worker_state["error"] = str(e)
print(f"❌ [worker] 启动失败: {e}")
return
last_heartbeat = 0.0
while not worker_state["stop"]:
try:
task_id, task_spec = task_queue.get(timeout=1)
except queue.Empty:
# 空闲时跑心跳
if time.monotonic() - last_heartbeat >= HEARTBEAT_INTERVAL:
run_heartbeat(ctx)
last_heartbeat = time.monotonic()
continue
state_store.update_task(task_id, state_store.TASK_RUNNING)
print(f">> [worker] 执行任务 #{task_id}: {task_spec}")
status, error = dispatch_task(ctx, task_spec)
state_store.update_task(task_id, status, error)
print(f">> [worker] 任务 #{task_id} 完成: {status} {error or ''}")
try:
ctx.stop()
except Exception:
pass
worker_state["ready"] = False
print(">> [worker] 已退出。")
def _enqueue_undelivered(site):
"""定时 job把该站 undelivered 任务投到队列worker 串行处理;本线程不碰 Playwright"""
try:
tid = state_store.create_task(site, "undelivered")
task_queue.put((tid, {"site": site, "kind": "undelivered"}))
print(f">> [定时] 投递 {site}/undelivered 任务 #{tid}")
except Exception as e:
print(f">> [定时] 投递 {site} 失败: {e}")
def _reschedule_site(site):
"""按持久化配置(重新)注册或取消该站的每日定时 job。"""
job_id = f"site_{site}"
try:
scheduler.remove_job(job_id)
except Exception:
pass
cfg = state_store.get_all_config().get(site)
if not cfg or not cfg.get("schedule_enabled") or not cfg.get("schedule_time"):
return
try:
hh, mm = cfg["schedule_time"].split(":")
scheduler.add_job(
_enqueue_undelivered,
CronTrigger(hour=int(hh), minute=int(mm)),
args=[site],
id=job_id,
replace_existing=True,
)
print(f">> [定时] 已注册 {site} 每日 {cfg['schedule_time']} 下载")
except Exception as e:
print(f">> [定时] 注册 {site} 失败: {e}")
@asynccontextmanager
async def lifespan(_app):
"""服务启停:起 worker 线程 / 通知 worker 停。"""
state_store.init_db() # 先建表/迁移状态库,确保早于 worker 就绪的 /api/status 可用
for site in ALL_SITES: # 按持久化配置注册各站每日定时 job
_reschedule_site(site)
scheduler.start()
print(">> [定时] 调度器已启动")
t = threading.Thread(target=_worker_loop, daemon=True)
worker_state["thread"] = t
t.start()
yield
scheduler.shutdown(wait=False)
print(">> [定时] 调度器已停止")
worker_state["stop"] = True
t.join(timeout=10)
app = FastAPI(title="InboundVerify 服务端", lifespan=lifespan)
class TaskRequest(BaseModel):
site: str
kind: str
@app.post("/tasks")
def create_task(req: TaskRequest):
"""提交任务 {site, kind} → 入队,返回 task_id。"""
# 【P0】后端未就绪时直接拒绝避免任务在 worker 启动前入队卡死
if not worker_state["ready"]:
raise HTTPException(status_code=409, detail="后端尚未就绪,请等待各站点登录完成后再操作")
if (req.site, req.kind) not in TASK_HANDLERS:
raise HTTPException(status_code=400, detail=f"无效任务: {req.site}/{req.kind}")
task_id = state_store.create_task(req.site, req.kind)
task_queue.put((task_id, {"site": req.site, "kind": req.kind}))
return {"task_id": task_id}
@app.get("/tasks/{task_id}")
def get_task(task_id: int):
t = state_store.get_task(task_id)
if not t:
raise HTTPException(status_code=404, detail="任务不存在")
return t
@app.get("/tasks")
def list_tasks(limit: int = 20):
return state_store.list_tasks(limit)
@app.get("/status")
def get_status():
"""各站登录态 + 数据态(前端状态盘用),另含 worker 就绪状态。"""
return {
"worker_ready": worker_state["ready"],
"worker_error": worker_state["error"],
"sites": state_store.get_all_status(),
}
class ConfigRequest(BaseModel):
expected_offset: Optional[int] = None
actual_offset: Optional[int] = None
schedule_enabled: Optional[bool] = None
schedule_time: Optional[str] = None
settings: Optional[Dict[str, str]] = None
def _default_cfg():
return {
"expected_offset": 0,
"actual_offset": 0,
"schedule_enabled": False,
"schedule_time": "",
}
@app.get("/config")
def get_config():
"""各站配置:应到/实到偏移 + 每日定时(百世偏移恒 0"""
cfg = state_store.get_all_config()
return {site: cfg.get(site, _default_cfg()) for site in ALL_SITES}
@app.put("/config/{site}")
def set_config(site: str, req: ConfigRequest):
if site not in ALL_SITES:
raise HTTPException(status_code=400, detail=f"未知站点: {site}")
cur = state_store.get_all_config().get(site, _default_cfg())
# 应到/实到偏移(百世锁定当天)
for kind, val in (
("expected", req.expected_offset),
("actual", req.actual_offset),
):
if val is not None:
if site == "百世":
raise HTTPException(
status_code=400, detail="百世固定下载当天,不可配置偏移"
)
state_store.set_offset(site, kind, val)
# 每日定时(所有站可配,含百世)
if req.schedule_enabled is not None or req.schedule_time is not None:
enabled = (
req.schedule_enabled
if req.schedule_enabled is not None
else cur["schedule_enabled"]
)
time_str = (
req.schedule_time if req.schedule_time is not None else cur["schedule_time"]
)
state_store.set_schedule(site, enabled, time_str)
_reschedule_site(site)
# 站点专属配置(密码/账号/路径…)
if req.settings:
for k, v in req.settings.items():
state_store.set_setting(site, k, v)
cfg = state_store.get_all_config().get(site, _default_cfg())
cfg["settings"] = state_store.get_site_settings(site)
return cfg
@app.get("/config/{site}/settings")
def get_site_settings_api(site: str):
"""单站专属配置(密码/账号/路径…;不进 5s 轮询)。"""
if site not in ALL_SITES:
raise HTTPException(status_code=400, detail=f"未知站点: {site}")
return state_store.get_site_settings(site)
REPORT_FILE = "应到未到数据.xlsx"
@app.get("/report")
def download_report():
"""下载 output/应到未到数据.xlsx跑比对生成未生成 404"""
path = os.path.join(OUTPUT_DIR, REPORT_FILE)
if not os.path.isfile(path):
raise HTTPException(status_code=404, detail="报告尚未生成,请先跑比对")
return FileResponse(path, filename=REPORT_FILE)
@app.get("/data/{filename}")
def download_data(filename: str):
"""下载 downloads/ 下的数据文件(防路径穿越)。"""
if not filename or "/" in filename or "\\" in filename or ".." in filename:
raise HTTPException(status_code=400, detail="非法文件名")
path = os.path.join(DOWNLOAD_DIR, filename)
# 双重校验:解析后绝对路径仍在 DOWNLOAD_DIR 内
if not os.path.abspath(path).startswith(os.path.abspath(DOWNLOAD_DIR) + os.sep):
raise HTTPException(status_code=400, detail="非法路径")
if not os.path.isfile(path):
raise HTTPException(status_code=404, detail="文件不存在")
return FileResponse(path, filename=filename)
if __name__ == "__main__":
uvicorn.run(app, host="0.0.0.0", port=8000)