docs: prevent double-decoding CVP IS links
Validate skill / validate (push) Successful in 4s

This commit is contained in:
2026-09-03 06:09:08 +00:00
parent 53923ee180
commit f0f3322a9a
2 changed files with 3 additions and 0 deletions
+2
View File
@@ -45,6 +45,8 @@ Live validation on 2026-08-19 used `resourceId=9279980`. All three routes return
7. Parse each HTML table row as a record. The notice **type** is commonly an anchor in the first cell, while the procurement **title** is plain text in the second cell—not an anchor. Therefore, anchor-text extraction alone can falsely report zero IT candidates even when a relevant title is visible. Extract and normalize all cells, keep the first-cell attachment URL, then keyword-filter the second-cell title.
8. Open promising notice-type links. They can be forced PDF attachments rather than HTML pages.
When a DOM/HTML parser has already decoded attribute entities, pass its `href` value directly to `urljoin`; do not apply `html.unescape()` again. A second generic unescape can reinterpret the beginning of the real `&noticeType` query key as the legacy `&not` entity and produce a corrupted `¬iceType` parameter. Before fetching a retained notice link, parse its query and require the expected keys (`resourceId`, `documentId`, `noticeType`, and `noticeId` when present). When working from raw HTML instead, decode `&` once and perform the same key check.
## Direct PDF behavior
A notice link commonly has this form: