Set-based port of scheduler.ts evaluateEquivalenceResearch(), applied once
live to every pending row with re_research_due_at IS NULL (the root cause
this pairs with a code fix for, previous commit). Result: 41,963 approved,
187,599 rejected as technical mismatch, 57,683 rejected as insufficient
evidence (re-research in 21 days), 2,436 rejected for no recent competitor
price. Backup of pre-triage state in
equivalence_pending_triage_backup_20260713.
catalog-reconcile.ts inserted new pending equivalences with
re_research_due_at always NULL. The daily maintenance:re-research-equivalences
job only evaluates rows where re_research_due_at IS NOT NULL AND <= NOW() -- so
the job that was built to auto-approve/reject the queue never saw fresh
pending rows. Backlog reached 289,681 rows before anyone noticed (found live
2026-07-13, one-time triage in sql/139: 41,963 approved, 247,718 rejected,
57,683 of those flagged for re-research once missing specs get enriched).
Fix: pending inserts now get re_research_due_at=NOW() so the job drains them
going forward. Also: reject branch now distinguishes data-gap rejects
(missing wavelength/fiber/reach evidence -> re-checked in 21 days once a
scraper fills the gap) from confirmed-mismatch/no-price/low-confidence
rejects (permanent, re_research_due_at stays NULL -- nothing changes on a
re-check against the same stored specs).
126 only guards writes to compatibility; a later transceivers.speed_gbps update
(e.g. daily enrich-modulation-fec.sh self-heal) never re-validated existing
compatible spec_match rows. Found live 2026-07-12: 189 rows had drifted
physically-impossible again. One-time remediation (189 rows -> incompatible,
backed up to compat_speed_drift_backup_20260712) + AFTER UPDATE OF speed_gbps
trigger on transceivers closes the gap for good. check_switch_quality() passes.
Add parseStandardName(title) to derive standard from product name
(SR, LR, ER, DR4, FR4, ZR+, etc.) since the V2 API has no protocol field.
Apply to all future syncs. Companion SQL backfill added 1,072 more
standard_names (total 2,281 / 29.8% FX coverage, up from 15.8%).
Phase 1 — DB: adds dwdm_channel, cwdm_channel, cdr_type, protocol_family
columns to transceivers (migration 134). Backfill from Flexoptix notes:
1,209 standard_names (+1,579%), 1,177 cdr_types, 2,498 wdm_types,
7,918 protocol_families.
Phase 2 — API sync: flexoptix-api-sync.ts now writes standard_name,
wdm_type, cdr_type, cdr_support, protocol_family on every sync.
Phase 3 — Scoring: catalog-reconcile.ts adds hard rejects for
protocol_family/temp_range/wdm_type mismatches, -25pt standard_name
mismatch penalty, CDR type scoring, and requires fx.standard_name to
be set before auto_approve — closes the 77k false-positive gap.
Public price-comparison and hot-topics routes previously had no filter
for the flexoptix/flexoptix-selfconfigure/flexoptix-msrp marketplaces.
Flexoptix API data is customer-confidential and must not leave the system.
Adds marketplace NOT LIKE flexoptix% filter to:
- /api/price-comparison (top-50 CTE)
- /api/price-comparison/summary (both CTEs)
- /api/price-comparison/:sku (per-vendor subquery)
- /api/hot-topics price_drops query (vendor name filter)
Parse Flexoptix API compat entries (compatible_to_vendor + original_part_number)
and batch-upsert into flexoptix_product_map during each catalog sync.
- Fix extractCompatibility() to recognise compatibleToVendor key from API
- Add batchUpsertProductMap(): 500-row batch INSERTs with ON CONFLICT update
- Add writeProductMapEntries() per product, called after price/stock writes
- Track mapWrites in FlexoptixSyncResult and log in completion message
First sync: 9254 upserts → 2903 unique rows, 95 vendors, 343 Flexoptix SKUs.
Example: CWDM-SFP10G-1330 (Cisco) → P.1696.10.D.I (EUR 150.77, SFP+ 10G).
- Add two-phase sync: bulk catalog fetch then targeted batch enrichment
(50 SKUs/call, ~10 extra calls) to retrieve self_configure_price field
which is absent from paginated responses
- Write self_configure_price as flexoptix-selfconfigure marketplace rows
in price_observations (343/346 products enriched on first run)
- Add list_price support (flexoptix-msrp marketplace) for future use
- Use API-provided type field (Transceiver/Cables/Accessories) in
categoryFor() instead of inferring from title alone
- Add productType field to CatalogProduct interface
Root cause of the equivalence review-queue backlog: spec-matcher read
transceivers.wavelengths (text), but the enricher wrote wavelength_tx_nm
(int) — 27,908 rows had the wavelength in the wrong column, and tx_nm
itself defaulted incorrectly to 1310nm in ~22% of disagreements with the
authoritative part-derived value.
sql/131: backfill + sync trigger from wavelength_tx_nm into wavelengths.
sql/132: correct authority (bare-number wavelengths, not tx_nm) + format
normalization + valid promotion on the full spec-matcher gate.
sql/133: one-directional reach-tier substitution (higher reach may satisfy
a lower-reach requirement, never the reverse; ZR/coherent excluded — real
link-budget/receiver-overload risk, not a simple tier comparison). Tagged
as its own status (compatible_superset), not merged into auto_approved.
Two bugs blocked all Flexoptix stock_observations:
1. db.ts: compatible_brands null param needs ::text[] cast; price_currency needs
::bpchar cast in stock_observations INSERT
2. flexoptix-api-sync.ts: UPDATE transceivers used $1 in WHERE IS NOT NULL clause
without explicit type cast — PostgreSQL cant infer type of untyped null params
in WHERE context. Fixed with $1::text / $2::text casts.
Result: 346 stock observations written on first run after fix.
upsertStockObservation dedup guard was preventing writes for unchanged
stock data (e.g. ATGBICS all-OOS for 55d), causing v_scraper_health to
mark sources as dead even when scrapers run normally. Skip dedup when
last observation is >7 days old so health view stays current.
- Replace plaintext tip (transceiver_db) and llm_gateway DB passwords with
env-var references guarded by fail-fast checks, across 15 scripts + sql runbook.
Covers quoted exports, unquoted psql calls, python env dicts and
env-or-hardcoded-default fallbacks.
- Untrack training-data/*.jsonl (~48MB, kept on disk) and gitignore *.jsonl.
- Add .security-scan-allowlist for legit false positives: vendor name flexoptix,
own homelab infra, sibling-project codenames in the sync journal, env-resolved
DB URIs, example placeholders, ML tokenizer tokens.
- Pre-push leak scanner goes from 397 to 0 findings, green without --no-verify.
Follow-up (not in this commit): rotate the leaked values (llm_gateway first,
tip fleet-wide coordinated); git history still holds the old passwords.
- escalation-sink: add NTFY_TOKEN Bearer auth header support
- self-heal-guards: pass max_load from GuardConfig to isLoadAcceptable
- control-loop: include MAX_LOAD in guardConfig so DQO_MAX_LOAD env is respected
The Adapter (sql/129) seeded all vendor sources as erik-safe, but TIP data is
fetched by the 3 residential Pi 5 workers (tashi/portia/ester), all now
Playwright-capable — not Erik (datacenter IP, blocked by FS.com/ATGBICS).
14 sources -> pi-fetch (fixes gbics/sfpcables never-dispatching under erik-safe).
Config only; loop stays in Plan-Mode. 10gtek stays proxmox-heavy.
Close the missing feedback loop. Detection (v_scraper_health, v_data_quality_alerts,
check_switch_quality) was accurate but inert; dispatch was blind fixed-cadence
(boss.send() is never called from any health path). tip-dqo is a separate, read-mostly
control plane that closes DETECT -> PRIORITIZE -> DISPATCH -> VERIFY -> RECHECK against
observation-table freshness (never pg-boss job state), with DB-enforced guardrails.
- sql/129: source_registry (signal->queue map + per-source SLA + guardrail contract +
global kill-switch), self_heal_dispatch append-only ledger, scraper_run_log,
acknowledged_exceptions, v_dqo_targets ranked view.
- orchestrator/: control-loop, self-heal-guards (pure, unit-tested), escalation-sink
(ntfy + no-op fallback), index entrypoint. Reuses the verification-robots dispatch
idiom (extracted shared dispatchQueues) and the quality views as-is.
- Guardrails: kill-switch, plan-mode default, per-source cooldown/budget/backoff/
circuit-breaker, load-guard, erik-safe profile cap, delta-gated escalation. Root-cause
routing escalates code_fix/blocked (e.g. ATGBICS silent stock-write break) to a human
instead of looping retries; only staleness is auto-dispatched.
- security: remove hardcoded 'tip_dev_2026' DB password fallback across 6 sites ->
fail-closed requireDbPassword() / dbConnectionString().
SELF_HEAL_PLAN_ONLY defaults true (dispatches nothing; writes plan rows for review).
Verified against the live DB: detects ATGBICS-stock dead/54d -> escalate, Flexoptix-stock
stale/30d -> plan-dispatch; guards + ledger correct; 8/8 guard unit tests pass.
- fn_strip_double_url_prefix + one-time fix for 436 rows whose scraped URLs had a
duplicated host prefix (https://host.comhttps://host.com/x) -> unfetchable.
Ingest guard trg_transceiver_url_normalize prevents recurrence.
- v_datasheet_enrichment_worklist: pipeline input listing every switch/transceiver
needing a field filled from source, with its best source URL + type.
The script hardcoded the pre-rotation POSTGRES_PASSWORD (retired 2026-06-04,
already invalid). Source the current password from ~/.tip/.env (git-ignored,
exported) instead — same pattern as the FS.com residential runner.
One article (The Next Platform) had published_at 2026-07-15 while scraped
2026-06-19 -- an article cannot be published after it was scraped. One-time
clamp to the scrape date + BEFORE INSERT/UPDATE trigger to prevent recurrence.
- trg_compat_speed_guard: BEFORE INSERT/UPDATE on compatibility forces
algorithmic (spec_match/community) rows whose transceiver is faster than the
switch max port speed to status=incompatible. Curated rows never touched.
Makes the "no physically-impossible compatibility" invariant structural, so
it survives a scraper regression.
- check_switch_quality(): on-demand proof; CI gate is
`select bool_and(pass) from check_switch_quality() where enforced` (must be
true). Data-source gaps (coding coverage, speed<=0, datasheet) tracked as
enforced=false so they do not fail the gate before the data exists.
Root cause: the form-factor fallback in flexoptix-compat matched transceivers by
form factor only, ignoring speed, so a 400G OSFP port matched every OSFP module
incl. 1.6T (5686+ impossible rows across 266 switches).
- flexoptix-compat: fallback now filters transceiver speed <= switch max port
speed (derived from ports_config + stored max_speed_gbps). Prevents regeneration.
- sql/125: evidence-based effective-ceiling (stored/ports/curated) views;
flag 5811 impossible spec_match rows status=incompatible (curated rows never
touched, understated switches never false-flagged); backfill max_speed_gbps for
327 switches from ports_config; re-derive speed for 306 speed<=0 transceivers
(form-factor capped, high-confidence only). v_switch_quality_alerts is the task
list; v_switch_understated_ports surfaces switches needing datasheet enrichment.
Reading packages/scraper showed QSFPTEK/ATGBICS populate quantity_available
100%; they simply lack FS.com-only units_sold/warehouse-split. The 123 alert
keyed on those FS.com fields and false-positived QSFPTEK. Rekey to
quantity_available OR warehouse_qty so the alert list drains to empty; surface
pct_qty_available/pct_vendor_ts in v_stock_completeness.
FS.com renders per-warehouse availability as a bare word-date on each line
("2 Stk. im DE-Lager, 8. Juli 2026 Lieferbar"). The old regex only matched a
"Lieferung:"/"erwartet:" prefix or numeric dd.mm.yyyy, so warehouse_de/global
delivery_date + backorder_estimated_date were 0% populated across 85k rows, and
lead_time_days was never written at all.
- new fs-com-parse.ts: pure, unit-tested parsers (13 node:test cases).
- extractDeliveryDate finds the bare word/numeric date per warehouse, bounded to
its own line so a date never bleeds to a neighbouring warehouse.
- computeLeadTimeDays = min(delivery_date - today), floored at 0; threaded into
upsertStockObservation (now writes lead_time_days) and upsertPriceObservation.
- parseGermanQty: K/M abbreviations now treat '.' as a decimal ("4.8K"=4800,
"1.4M"=1.4M) instead of stripping it (was "1.4M"->14,000,000 — the outlier).
- units_sold sanity gate at ingest (FS_MAX_UNITS_SOLD, default 2,000,000).
- sql/122: purge impossible historical units_sold, re-document columns as collected.
sql/119: product_type enum (module|dac|aoc|aec|switch|nic|accessory) computed by
trigger trg_transceivers_biu from part_number/category/reach on every insert/update.
Backfilled 41008 rows (36284 module, 2791 dac, 1835 aoc, 29 aec, 27 accessory,
26 switch, 16 nic). Cables and switches/NICs no longer counted as modules;
BASE-T/RJ45 copper modules protected from cable demotion; AOC/DAC keyword beats
the optical-token guard. Consumers filter product_type=module.
sql/120: conservative form_factor recovery. Corrects only provably-wrong rows
(form_factors.max_speed < speed_gbps) via an explicit title token; 87 fixed
(e.g. 400G cables mislabelled SFP+ to QSFP-DD), 947 opaque rows left untouched
(never guessed). Scraped original preserved in new column form_factor_raw.
sql/121: document lead_time_days as not-collected (0 percent populated on all
time-series tables) so no consumer fabricates a default.
scraper/utils/db.ts: note that product_type/form_factor are trigger-maintained.
A switch presents transceiver compatibility ONLY if ports_verified=true (its port
config confirmed against the manufacturer datasheet). Unverified switches return
NO suggestions rather than potentially-wrong data. New columns: ports_verified,
ports_verified_source, ports_verified_at. Verified so far against datasheets:
HPE Aruba CX 10000-48Y6C, NVIDIA SN3700 (corrected 100G->200G QSFP56), SN4700,
SN5400, SN5600. 262 switches still pending datasheet verification.
Chassis switches use ports_config keys like 'max_per_slot_100G_QSFP28' (speed in
the middle), not just '100G_QSFP28'. The '^[0-9.]+G' anchor failed -> port_speed
NULL -> 0 suggestions (e.g. Cisco 8818 router). Now extracts the speed token
wherever it sits before the cage. Cisco 8818: 0 -> 55, all Cisco-codeable, max 400G.
Physical fit alone is not real compatibility — a module that seats in the cage
still won't run unless Flexoptix can code it for the switch's platform. Added a
coding requirement: each suggestion must have an fx_compatibilities entry for the
switch's mapped vendor (HPE/Aruba->aruba, Cisco->cisco, NVIDIA->mellanox, etc.;
whitebox/unmapped -> MSA Standard fallback).
CX 10000-48Y6C: 629 physically-fitting -> 283 actually-codeable-for-Aruba. All
283 verified to carry an Aruba coding. Every suggestion is now catalog-confirmed,
datasheet-accurate, physically fitting AND guaranteed codeable to run in the switch.
nominalSpeedFromSpecs missed '1.6T Ethernet' protocol entries (only matched 'G'),
returning 800 for twin-port 1.6T parts instead of 1600. Added Terabit handling +
guarded against 'Nx 800G' aggregate partial-matches. After re-deriving all 805
catalog-confirmed FX parts from their stored datasheet specs, DB now matches the
Flexoptix datasheet 100%: Form Factor 802/802, Speed 801/801, 0 deviations.
~107 parts were misattributed to the Flexoptix vendor but do not exist in the
live Flexoptix catalog (FX uses dot-SKUs like S.1606.28.KD; these use dash-SKUs
like FOT-AX-200G). They polluted 'BEI FLEXOPTIX BESTELLEN' with non-orderable
phantom parts carrying name-guessed specs. Now require fx_specifications IS NOT
NULL — every suggestion is confirmed in Flexoptix's API with authoritative
datasheet specs. CX 10000-48Y6C: 801 -> 630 suggestions, all catalog-confirmed,
0 physically impossible, 0 phantoms.
Root cause of unreliable speeds: the bulk Flexoptix products API returns NO speed
or form-factor field, so the catalog sync guessed them by parsing the product
NAME (e.g. 'FO-109010-CWDM' -> 100000G). Cryptic FX codes like 'O.164HG2.2.C:Sx'
are unparseable, producing garbage.
The detail enricher already pulls per-SKU specifications=1 (rate-limited) but only
wrote secondary fields. Now it also derives:
form_factor <- structured 'Form Factor' spec (authoritative datasheet value)
speed_gbps <- highest Ethernet rate in 'Supported Protocols', fallback to the
'Bandwidth' line-rate upper bound mapped to nominal
Both OVERWRITE the corrupt bulk values (COALESCE(spec, existing)). Never derived
from the product name. Verified: 100/100 freshly-enriched FX parts now have
physically-consistent form_factor/speed (0 contradictions), incl. uncrackable
codes correctly resolved to OSFP/800G, QSFP-DD800/800G etc.
getFlexoptixSuggestions matched ONLY by form factor, discarding the speed encoded
in each ports_config key (e.g. '100G_QSFP28'). Corrupt transceiver speed_gbps
values (400G/200G/128G/100000G mislabeled as QSFP28) leaked through, so a 100G
switch showed impossible '400G QSFP28' / '100T QSFP28' suggestions.
Now parses (speed, form_factor) from each port key and requires every suggested
module to (a) mechanically seat in the cage — precise port-FF -> accepted-module-FF
map, a QSFP28 cage takes QSFP+/28/56 but never QSFP-DD — and (b) have
speed_gbps <= the port's speed. CX 10000-48Y6C (25G SFP28 + 100G QSFP28) now
returns only valid <=25G SFP / <=100G QSFP modules; 0 physically impossible
entries (was 4 garbage groups). Belt-and-suspenders: even with corrupt speed data,
nothing oversized can reach a customer-facing suggestion.