This commit is contained in:
@@ -45,6 +45,8 @@ Live validation on 2026-08-19 used `resourceId=9279980`. All three routes return
|
||||
7. Parse each HTML table row as a record. The notice **type** is commonly an anchor in the first cell, while the procurement **title** is plain text in the second cell—not an anchor. Therefore, anchor-text extraction alone can falsely report zero IT candidates even when a relevant title is visible. Extract and normalize all cells, keep the first-cell attachment URL, then keyword-filter the second-cell title.
|
||||
8. Open promising notice-type links. They can be forced PDF attachments rather than HTML pages.
|
||||
|
||||
When a DOM/HTML parser has already decoded attribute entities, pass its `href` value directly to `urljoin`; do not apply `html.unescape()` again. A second generic unescape can reinterpret the beginning of the real `¬iceType` query key as the legacy `¬` entity and produce a corrupted `¬iceType` parameter. Before fetching a retained notice link, parse its query and require the expected keys (`resourceId`, `documentId`, `noticeType`, and `noticeId` when present). When working from raw HTML instead, decode `&` once and perform the same key check.
|
||||
|
||||
## Direct PDF behavior
|
||||
|
||||
A notice link commonly has this form:
|
||||
|
||||
Reference in New Issue
Block a user