`.agents/setup` now pre-loads the CLI's local data after the build: it
installs the published ownership snapshot (`idx ownership sync`, the
repo's preferred bootstrap path) and warms the file cache for the shipped
`stocks` read commands, so a fresh orb can query ownership data and serve
`--offline` reads immediately.
Setup also adopts the dev shell environment for its remaining steps via
`nix print-dev-env` (the same mechanism the login-shell hook uses), and
the orb guidance block is replaced in place so text updates reach orbs
that already have one without accumulating blank lines.
Amp-Thread-ID: https://ampcode.com/threads/T-01a0ba76-221d-728b-9d68-45ab48750d3d
Co-authored-by: The Diligence Dev <thediligencedev@gmail.com>
Fresh Amp orbs start without a Rust toolchain or Nix, so every agent would
repeat the same install. `.agents/setup` installs Nix (single-user),
realizes the flake dev shell, activates it for the login shells Amp and orb
services start, and warms target/ so the first agent command is incremental.
`.agents/resume` re-checks the toolchain after a pause.
Also ignore .amp/portals/ so orb runtime files stay out of git status.
Amp-Thread-ID: https://ampcode.com/threads/T-01a0ba76-221d-728b-9d68-45ab48750d3d
Co-authored-by: The Diligence Dev <thediligencedev@gmail.com>
- sanitize_current_ratio: and_then if/else -> filter
- screener/graph sorts: sort_by with reversed cmp -> sort_by_key(Reverse)
- insert_ksei_holdings_rows: manual Result match -> ?
Fixes the Clippy step on ubuntu-latest (rust 1.97) that the nix
toolchain (1.93.1) did not flag.
Newer KSEI above1 PDFs use spaces in column headers (SHARE CODE,
INVESTOR NAME, INVESTOR CLASSIFICATION) instead of the underscored
form (SHARE_CODE, INVESTOR_NAME, INVESTOR_TYPE) that the parser
schema markers expected. This caused valid holder-register PDFs to
be misclassified as LegacyAboveFivePercent (false positive on
REKENING TAMPUNGAN KSEI text in the data), blocking import.
Changes:
- Add normalize_stext_for_classification() that replaces spaces with
underscores inside text="..." attributes before schema marker
matching, so both old and new header layouts are recognized
- Add INVESTOR_CLASSIFICATION to HOLDER_REGISTER_SCHEMA_MARKERS and
HEADER_LABELS to match the renamed column
- Update ANNOUNCEMENT_WRAPPER and ABOVE_FIVE marker constants to use
underscores (matching the normalized text)
- Add test fixture and test case for the spaced-header layout
- Bump version to 0.2.3
Tested against the live June 2026 KSEI PDF (as_of 2026-05-29):
idx ownership import --url <latest-pdf-url>
-> Imported 7124 rows for 956 tickers (as of 2026-05-29)
The April 2026 IDX ownership PDF omits the zero-valued scrip column for
some holders, causing the parser to leave holdings_scripless/scrip as
empty strings. This broke normalize_ksei_row which called parse_id_number
on empty input.
Two-layer fix:
- Parser: pop_numeric_tail handles 2-column (missing scrip) rows by
defaulting scrip to "0"
- Normalizer: parse_id_number_or_zero backstop treats empty component
share fields as 0 while still requiring total_shares
Also includes AGENTS.md refresh, formatting cleanup (rustfmt), and
version bump to 0.2.2.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add a compact parser fixture excerpted from the real 2026-03-10 IDX above1 lamp1 mutool stext output and assert the first live row still parses as expected.
This keeps the post-merge ownership parser coverage grounded in the supported above1 layout instead of relying only on synthetic line fixtures.
Clear the smoke cache before each cache-group warm case so the following offline and stale-cache checks always start from a fresh baseline instead of earlier group state.
Document the runner behavior in the smoke guide to make the cache-group sequencing explicit.
Refactor KSEI PDF parsing for the live above-1 holder-register layout, add IDX announcement discovery plus browser-impersonated PDF fetches, and update the ownership docs to mark Batch 1 complete and scope Batch 2 around above1 hardening and unsupported legacy inputs.