The 14 false curated rows had been raising the evidence-based ceiling for
those 1G-only switches, masking 261 algorithmic spec_match rows claiming >1G
transceivers on the same hardware. With the ceiling honest again,
v_compat_speed_errors correctly flagged them; standard 125-B/138 remediation
applied (backup in unmasked_spec_match_backup_20260718). Gate green again.
Manual datasheet-verified correction, the one class the automated guards
(126/138/141) deliberately never touch: curated vendor_compat rows claiming
10G Flexoptix transceivers work in 1G-only switches. Verified against Cisco
primary docs 2026-07-18: ASR9000 transceiver support matrix (C78-624747-06)
shows A9K-40GE-L/B/E accept SFP only (XFP/SFP+/QSFP/CFP all No), and the
Catalyst 9300 datasheet specifies C9300-48S/24S as 1G SFP fixed downlinks
(10G only via separate NM uplink modules, tracked as own switch entities).
C9400-LC-48S/24S + C9600-LC-48S are 1G SFP line cards. Backup in
curated_compat_10g_on_1g_backup_20260718. v_switch_understated_ports now 0.
Drafts CHANGELOG_PENDING.md entries from Conventional Commit subjects on
every push to main -- feat/fix/refactor/perf/security bucketed into
Added/Fixed/Changed, chore/ci/test/docs/style/merge skipped, anything
mentioning internal infra hostnames dropped. No LLM call, no network
access. Never touches CHANGELOG.md -- still needs a human to fold an
entry in. Rolled out fleet-wide from the PeerCortex pilot (2026-07-18).
Investigated the coding-matrix gap per Rene: vendor_compat is NOT derived
per-SKU from fx_compatibilities (found rows with vendor_compat fully populated
but fx_compatibilities empty, and vice versa -- zero correlation). It is a
shared OEM-coding-pattern reference template keyed by (speed_gbps,
form_factor): only 11 distinct templates exist across the whole Flexoptix
catalog, essentially 1:1 with family (one ambiguous exception, 100G/SFP,
deliberately excluded). Backfilled 2,587 rows from sibling templates (fill-only,
skips any family with 0 or >1 distinct templates). Flexoptix >=100G coding gap
dropped from 192 to 25 missing. New trigger (trg_fill_vendor_compat_from_family)
keeps this self-healing: every future Flexoptix product inherits its family
template automatically once speed_gbps+form_factor are known (after detail
sync), instead of needing another one-off backfill.
Deliberately NOT extended to competitor products in this pass -- needs a
separate, more careful evaluation of what this data actually represents for
non-Flexoptix parts before propagating via equivalence.
Re-run of sql/125 Remediation C against the current v_transceiver_speed_errors
gap (596 as of 2026-07-17, mostly Cisco Systems 230 + Juniper Networks 220).
Same high-confidence-only logic: part-number speed token, capped at the form
factor physical max, no blind updates. 339/596 recovered mechanically, 257
genuinely need datasheet work (GE/FE/OC-rate parts with no G-token). Verified
the new switch-side (141) and transceiver-side (138) drift guards handled the
339 speed changes correctly: v_compat_speed_errors stayed at 0 throughout.
Preventive companion to 126 (compatibility writes) + 138 (transceiver speed
changes): the same physically-impossible-compat invariant also depends on
switches.max_speed_gbps/ports_config via v_switch_effective_max, but nothing
guarded THAT write direction. Audit 2026-07-17: currently a no-op live
(nothing updates those columns today outside the one-time sql/125 backfill),
but the documented open item -- 110 switches need ports_config datasheet
enrichment -- is exactly the future job that would trigger this gap the same
way the daily modulation/FEC cron triggered 138. Closing it now. Verified with
a synthetic in-transaction test (1,250 rows correctly flagged, then rolled
back).
One-time companion to commit e7299ac7 (spec-matcher.ts now inserts pending,
not auto_approved). Re-scores the pre-existing spec-basis auto_approved rows
against the rigorous evaluateEquivalenceResearch() criteria: 972 (67%)
genuinely hold up and stay auto_approved, 483 (33%) get demoted -- some
permanently (technical mismatch / no recent competitor price), some flagged
for re-research in 21 days (missing wavelength/fiber/reach evidence, not a
wrong match). Backup in spec_matcher_reclassify_backup_20260717.
spec-matcher.ts inserted every JOIN match straight into status=auto_approved
at a flat confidence=0.85 -- no standard_name check, no fiber_type check, no
recent-competitor-price check, none of the discrimination catalog-reconcile.ts
uses before auto-approving. That is real ground live: of the 1,455 currently
auto_approved spec-basis equivalences, ~19% would fail a recent-price check
and ~8% have a conflicting standard_name if re-scored against the rigorous
bar (see sql/140 for the one-time reclassification of the existing rows).
Fix: spec matches now insert as pending with re_research_due_at=NOW(), so
they flow through the same maintenance:re-research-equivalences job (fixed
2026-07-17, sql/139) that already scores catalog-reconcile.ts pending rows
-- real evidence-based confidence before anything is served as approved, not
a blanket 0.85. opn-matcher.ts is untouched: OPN matches are
manufacturer-confirmed ground truth from fx_compatibilities, not a heuristic
guess, so blanket auto_approve there remains correct.
Set-based port of scheduler.ts evaluateEquivalenceResearch(), applied once
live to every pending row with re_research_due_at IS NULL (the root cause
this pairs with a code fix for, previous commit). Result: 41,963 approved,
187,599 rejected as technical mismatch, 57,683 rejected as insufficient
evidence (re-research in 21 days), 2,436 rejected for no recent competitor
price. Backup of pre-triage state in
equivalence_pending_triage_backup_20260713.
catalog-reconcile.ts inserted new pending equivalences with
re_research_due_at always NULL. The daily maintenance:re-research-equivalences
job only evaluates rows where re_research_due_at IS NOT NULL AND <= NOW() -- so
the job that was built to auto-approve/reject the queue never saw fresh
pending rows. Backlog reached 289,681 rows before anyone noticed (found live
2026-07-13, one-time triage in sql/139: 41,963 approved, 247,718 rejected,
57,683 of those flagged for re-research once missing specs get enriched).
Fix: pending inserts now get re_research_due_at=NOW() so the job drains them
going forward. Also: reject branch now distinguishes data-gap rejects
(missing wavelength/fiber/reach evidence -> re-checked in 21 days once a
scraper fills the gap) from confirmed-mismatch/no-price/low-confidence
rejects (permanent, re_research_due_at stays NULL -- nothing changes on a
re-check against the same stored specs).
126 only guards writes to compatibility; a later transceivers.speed_gbps update
(e.g. daily enrich-modulation-fec.sh self-heal) never re-validated existing
compatible spec_match rows. Found live 2026-07-12: 189 rows had drifted
physically-impossible again. One-time remediation (189 rows -> incompatible,
backed up to compat_speed_drift_backup_20260712) + AFTER UPDATE OF speed_gbps
trigger on transceivers closes the gap for good. check_switch_quality() passes.
Add parseStandardName(title) to derive standard from product name
(SR, LR, ER, DR4, FR4, ZR+, etc.) since the V2 API has no protocol field.
Apply to all future syncs. Companion SQL backfill added 1,072 more
standard_names (total 2,281 / 29.8% FX coverage, up from 15.8%).
Phase 1 — DB: adds dwdm_channel, cwdm_channel, cdr_type, protocol_family
columns to transceivers (migration 134). Backfill from Flexoptix notes:
1,209 standard_names (+1,579%), 1,177 cdr_types, 2,498 wdm_types,
7,918 protocol_families.
Phase 2 — API sync: flexoptix-api-sync.ts now writes standard_name,
wdm_type, cdr_type, cdr_support, protocol_family on every sync.
Phase 3 — Scoring: catalog-reconcile.ts adds hard rejects for
protocol_family/temp_range/wdm_type mismatches, -25pt standard_name
mismatch penalty, CDR type scoring, and requires fx.standard_name to
be set before auto_approve — closes the 77k false-positive gap.
Public price-comparison and hot-topics routes previously had no filter
for the flexoptix/flexoptix-selfconfigure/flexoptix-msrp marketplaces.
Flexoptix API data is customer-confidential and must not leave the system.
Adds marketplace NOT LIKE flexoptix% filter to:
- /api/price-comparison (top-50 CTE)
- /api/price-comparison/summary (both CTEs)
- /api/price-comparison/:sku (per-vendor subquery)
- /api/hot-topics price_drops query (vendor name filter)
Parse Flexoptix API compat entries (compatible_to_vendor + original_part_number)
and batch-upsert into flexoptix_product_map during each catalog sync.
- Fix extractCompatibility() to recognise compatibleToVendor key from API
- Add batchUpsertProductMap(): 500-row batch INSERTs with ON CONFLICT update
- Add writeProductMapEntries() per product, called after price/stock writes
- Track mapWrites in FlexoptixSyncResult and log in completion message
First sync: 9254 upserts → 2903 unique rows, 95 vendors, 343 Flexoptix SKUs.
Example: CWDM-SFP10G-1330 (Cisco) → P.1696.10.D.I (EUR 150.77, SFP+ 10G).
- Add two-phase sync: bulk catalog fetch then targeted batch enrichment
(50 SKUs/call, ~10 extra calls) to retrieve self_configure_price field
which is absent from paginated responses
- Write self_configure_price as flexoptix-selfconfigure marketplace rows
in price_observations (343/346 products enriched on first run)
- Add list_price support (flexoptix-msrp marketplace) for future use
- Use API-provided type field (Transceiver/Cables/Accessories) in
categoryFor() instead of inferring from title alone
- Add productType field to CatalogProduct interface
Root cause of the equivalence review-queue backlog: spec-matcher read
transceivers.wavelengths (text), but the enricher wrote wavelength_tx_nm
(int) — 27,908 rows had the wavelength in the wrong column, and tx_nm
itself defaulted incorrectly to 1310nm in ~22% of disagreements with the
authoritative part-derived value.
sql/131: backfill + sync trigger from wavelength_tx_nm into wavelengths.
sql/132: correct authority (bare-number wavelengths, not tx_nm) + format
normalization + valid promotion on the full spec-matcher gate.
sql/133: one-directional reach-tier substitution (higher reach may satisfy
a lower-reach requirement, never the reverse; ZR/coherent excluded — real
link-budget/receiver-overload risk, not a simple tier comparison). Tagged
as its own status (compatible_superset), not merged into auto_approved.
Two bugs blocked all Flexoptix stock_observations:
1. db.ts: compatible_brands null param needs ::text[] cast; price_currency needs
::bpchar cast in stock_observations INSERT
2. flexoptix-api-sync.ts: UPDATE transceivers used $1 in WHERE IS NOT NULL clause
without explicit type cast — PostgreSQL cant infer type of untyped null params
in WHERE context. Fixed with $1::text / $2::text casts.
Result: 346 stock observations written on first run after fix.
upsertStockObservation dedup guard was preventing writes for unchanged
stock data (e.g. ATGBICS all-OOS for 55d), causing v_scraper_health to
mark sources as dead even when scrapers run normally. Skip dedup when
last observation is >7 days old so health view stays current.
- Replace plaintext tip (transceiver_db) and llm_gateway DB passwords with
env-var references guarded by fail-fast checks, across 15 scripts + sql runbook.
Covers quoted exports, unquoted psql calls, python env dicts and
env-or-hardcoded-default fallbacks.
- Untrack training-data/*.jsonl (~48MB, kept on disk) and gitignore *.jsonl.
- Add .security-scan-allowlist for legit false positives: vendor name flexoptix,
own homelab infra, sibling-project codenames in the sync journal, env-resolved
DB URIs, example placeholders, ML tokenizer tokens.
- Pre-push leak scanner goes from 397 to 0 findings, green without --no-verify.
Follow-up (not in this commit): rotate the leaked values (llm_gateway first,
tip fleet-wide coordinated); git history still holds the old passwords.
- escalation-sink: add NTFY_TOKEN Bearer auth header support
- self-heal-guards: pass max_load from GuardConfig to isLoadAcceptable
- control-loop: include MAX_LOAD in guardConfig so DQO_MAX_LOAD env is respected
The Adapter (sql/129) seeded all vendor sources as erik-safe, but TIP data is
fetched by the 3 residential Pi 5 workers (tashi/portia/ester), all now
Playwright-capable — not Erik (datacenter IP, blocked by FS.com/ATGBICS).
14 sources -> pi-fetch (fixes gbics/sfpcables never-dispatching under erik-safe).
Config only; loop stays in Plan-Mode. 10gtek stays proxmox-heavy.
Close the missing feedback loop. Detection (v_scraper_health, v_data_quality_alerts,
check_switch_quality) was accurate but inert; dispatch was blind fixed-cadence
(boss.send() is never called from any health path). tip-dqo is a separate, read-mostly
control plane that closes DETECT -> PRIORITIZE -> DISPATCH -> VERIFY -> RECHECK against
observation-table freshness (never pg-boss job state), with DB-enforced guardrails.
- sql/129: source_registry (signal->queue map + per-source SLA + guardrail contract +
global kill-switch), self_heal_dispatch append-only ledger, scraper_run_log,
acknowledged_exceptions, v_dqo_targets ranked view.
- orchestrator/: control-loop, self-heal-guards (pure, unit-tested), escalation-sink
(ntfy + no-op fallback), index entrypoint. Reuses the verification-robots dispatch
idiom (extracted shared dispatchQueues) and the quality views as-is.
- Guardrails: kill-switch, plan-mode default, per-source cooldown/budget/backoff/
circuit-breaker, load-guard, erik-safe profile cap, delta-gated escalation. Root-cause
routing escalates code_fix/blocked (e.g. ATGBICS silent stock-write break) to a human
instead of looping retries; only staleness is auto-dispatched.
- security: remove hardcoded 'tip_dev_2026' DB password fallback across 6 sites ->
fail-closed requireDbPassword() / dbConnectionString().
SELF_HEAL_PLAN_ONLY defaults true (dispatches nothing; writes plan rows for review).
Verified against the live DB: detects ATGBICS-stock dead/54d -> escalate, Flexoptix-stock
stale/30d -> plan-dispatch; guards + ledger correct; 8/8 guard unit tests pass.
- fn_strip_double_url_prefix + one-time fix for 436 rows whose scraped URLs had a
duplicated host prefix (https://host.comhttps://host.com/x) -> unfetchable.
Ingest guard trg_transceiver_url_normalize prevents recurrence.
- v_datasheet_enrichment_worklist: pipeline input listing every switch/transceiver
needing a field filled from source, with its best source URL + type.
The script hardcoded the pre-rotation POSTGRES_PASSWORD (retired 2026-06-04,
already invalid). Source the current password from ~/.tip/.env (git-ignored,
exported) instead — same pattern as the FS.com residential runner.
One article (The Next Platform) had published_at 2026-07-15 while scraped
2026-06-19 -- an article cannot be published after it was scraped. One-time
clamp to the scrape date + BEFORE INSERT/UPDATE trigger to prevent recurrence.
- trg_compat_speed_guard: BEFORE INSERT/UPDATE on compatibility forces
algorithmic (spec_match/community) rows whose transceiver is faster than the
switch max port speed to status=incompatible. Curated rows never touched.
Makes the "no physically-impossible compatibility" invariant structural, so
it survives a scraper regression.
- check_switch_quality(): on-demand proof; CI gate is
`select bool_and(pass) from check_switch_quality() where enforced` (must be
true). Data-source gaps (coding coverage, speed<=0, datasheet) tracked as
enforced=false so they do not fail the gate before the data exists.
Root cause: the form-factor fallback in flexoptix-compat matched transceivers by
form factor only, ignoring speed, so a 400G OSFP port matched every OSFP module
incl. 1.6T (5686+ impossible rows across 266 switches).
- flexoptix-compat: fallback now filters transceiver speed <= switch max port
speed (derived from ports_config + stored max_speed_gbps). Prevents regeneration.
- sql/125: evidence-based effective-ceiling (stored/ports/curated) views;
flag 5811 impossible spec_match rows status=incompatible (curated rows never
touched, understated switches never false-flagged); backfill max_speed_gbps for
327 switches from ports_config; re-derive speed for 306 speed<=0 transceivers
(form-factor capped, high-confidence only). v_switch_quality_alerts is the task
list; v_switch_understated_ports surfaces switches needing datasheet enrichment.
Reading packages/scraper showed QSFPTEK/ATGBICS populate quantity_available
100%; they simply lack FS.com-only units_sold/warehouse-split. The 123 alert
keyed on those FS.com fields and false-positived QSFPTEK. Rekey to
quantity_available OR warehouse_qty so the alert list drains to empty; surface
pct_qty_available/pct_vendor_ts in v_stock_completeness.
FS.com renders per-warehouse availability as a bare word-date on each line
("2 Stk. im DE-Lager, 8. Juli 2026 Lieferbar"). The old regex only matched a
"Lieferung:"/"erwartet:" prefix or numeric dd.mm.yyyy, so warehouse_de/global
delivery_date + backorder_estimated_date were 0% populated across 85k rows, and
lead_time_days was never written at all.
- new fs-com-parse.ts: pure, unit-tested parsers (13 node:test cases).
- extractDeliveryDate finds the bare word/numeric date per warehouse, bounded to
its own line so a date never bleeds to a neighbouring warehouse.
- computeLeadTimeDays = min(delivery_date - today), floored at 0; threaded into
upsertStockObservation (now writes lead_time_days) and upsertPriceObservation.
- parseGermanQty: K/M abbreviations now treat '.' as a decimal ("4.8K"=4800,
"1.4M"=1.4M) instead of stripping it (was "1.4M"->14,000,000 — the outlier).
- units_sold sanity gate at ingest (FS_MAX_UNITS_SOLD, default 2,000,000).
- sql/122: purge impossible historical units_sold, re-document columns as collected.
sql/119: product_type enum (module|dac|aoc|aec|switch|nic|accessory) computed by
trigger trg_transceivers_biu from part_number/category/reach on every insert/update.
Backfilled 41008 rows (36284 module, 2791 dac, 1835 aoc, 29 aec, 27 accessory,
26 switch, 16 nic). Cables and switches/NICs no longer counted as modules;
BASE-T/RJ45 copper modules protected from cable demotion; AOC/DAC keyword beats
the optical-token guard. Consumers filter product_type=module.
sql/120: conservative form_factor recovery. Corrects only provably-wrong rows
(form_factors.max_speed < speed_gbps) via an explicit title token; 87 fixed
(e.g. 400G cables mislabelled SFP+ to QSFP-DD), 947 opaque rows left untouched
(never guessed). Scraped original preserved in new column form_factor_raw.
sql/121: document lead_time_days as not-collected (0 percent populated on all
time-series tables) so no consumer fabricates a default.
scraper/utils/db.ts: note that product_type/form_factor are trigger-maintained.
A switch presents transceiver compatibility ONLY if ports_verified=true (its port
config confirmed against the manufacturer datasheet). Unverified switches return
NO suggestions rather than potentially-wrong data. New columns: ports_verified,
ports_verified_source, ports_verified_at. Verified so far against datasheets:
HPE Aruba CX 10000-48Y6C, NVIDIA SN3700 (corrected 100G->200G QSFP56), SN4700,
SN5400, SN5600. 262 switches still pending datasheet verification.