
Fig. 1ScholarTrack listing 43 fully-funded programmes. The filter rail states what each eligibility filter does and does not include; cards show coverage, verification date, and where funding is not itemised rather than leaving a blank.
A zero-cost dashboard for finding, filtering and tracking fully-funded international scholarships. Programme details live in structured JSONB columns rather than prose, so questions a wall of text cannot answer become filters — and every extracted value is verified against a verbatim quote from the funder's own page before it is allowed into the database.
IThe problem
Scholarship information is scattered across funder websites in prose, and aggregators mostly copy each other's summaries. The result is a search experience that cannot answer the questions applicants actually have — "scholarships in Japan that need no work experience and include a funded language year" is a reading exercise across forty tabs, not a filter.
The obvious fix is to run a language model over each page and store what it returns. The obvious fix is also how you end up telling someone the wrong deadline. An earlier version of this reconstructed deadlines from each programme's typical cycle and got Chevening wrong by a month. A scholarship deadline is a date somebody plans a year of their life around.
IIHow it works
Grounding — the control everything else rests on
Every value the extractor produces must carry a verbatim quote, and that quote is checked by string search against the fetched page. Values whose quote is not found in the source are discarded. Fabrication is caught mechanically, not by a human reading output and hoping.
Numbers must appear inside their own quote: a grounded "includes a monthly stipend" does not license an invented amount printed beside it. And a date more than one cycle old is refused even when its quote is perfectly real, because programmes run annually and a deadline read off a page the funder left up since 2019 is not this cycle's.
Absent is not false
An unstated requirement renders as "Not stated", never "Not required" — and the filters hold the same line. "No work experience required" returns only programmes whose pages say so. A row nobody has read is not evidence of anything, and quietly treating missing data as a negative would turn silence into a claim.
Dates are nullable throughout. A programme with no published deadline renders as "Dates not published" rather than borrowing last year's. The rule the whole design follows is that a blank field beats a plausible wrong one.
Why a field is blank
The run log could always say "12 of 25 fields filled". It could say nothing about the other thirteen, and those thirteen are not one thing. A blank had four causes fixed in four different files, and only one of them was the pipeline working correctly — so "improve coverage" was an unfalsifiable goal, and a crawl that never reached the benefits page looked exactly like a funder that publishes no stipend.
Every empty slot is now classified: never-crawled, not-retrieved, model-silent, grounding-rejected, or not-published. The order runs from the most specific evidence to the least, and not-published is the residual — the claim the system is least entitled to make, and the only one allowed to reach a student. The interface says "not stated by the programme" only when the pages on that subject were actually crawled and read.
Second sources fill the gaps, not just check the dates
Cross-verification originally existed to reconcile deadlines across sources. It could not fill anything, because the crawler is same-origin by construction: coverage, eligibility and documents were read exclusively from the funder's own site, so whatever a funder chose not to publish stayed blank permanently.
Second sources now fill the profile — gaps only, never overwriting, grounded against each source's own pages, and every filled field attributed to where it came from. Where a whole topic was starved rather than merely missed, a further pass re-crawls with quotas on the starved topics and links scored on the missing facts' own vocabulary, excluding pages already read.
Provenance is never defaulted
last_verified_at stays NULL until a human actually checks, and the interface says "Not verified" out loud rather than hiding the gap. "Verified" is the one status no crawler may write — it comes from a person, through a dedicated review route. Nothing the pipeline produces is visible in the app until a human applies it.
Two programs, one database
The web app reads scholarships and writes only a user's own tracking rows; it never holds a privileged key. The pipeline crawls funder sites and writes proposals into staging tables using the service-role key — which bypasses every RLS policy, so it is confined to scripts, never prefixed NEXT_PUBLIC_, and never imported by anything under src/. Both halves now deploy to Vercel.
A weekly Vercel cron re-reads what has gone stale. The same two crawlers can also be driven from an admin console in the browser, with an argument whitelist and a tick runner that lets a long crawl survive a serverless function ceiling.
Ghost accounts, so tracking works on the first click
A proxy runs on every request: it refreshes the Supabase session and, if there is none, calls signInAnonymously(). The visitor holds a real JWT and a stable UUID before the first Server Component renders, so "Track this scholarship" works immediately with no login wall.
Anonymous users hold real JWTs, so RLS covers them under the authenticated role exactly like email users — there is no second code path for "logged out". Linking an email later preserves the UUID, so every tracked scholarship and ticked checkbox survives the upgrade. Sign-out is deliberately not offered to anonymous users, because their tracker is only reachable through that session.
IIIDecisions
JSONB columns instead of prose descriptions
Coverage, eligibility, documents and application windows are stored as structured shapes rather than paragraphs. This is what makes the sidebar able to answer compound questions, and it is also what makes the "not stated" distinction expressible at all — prose has no way to be explicitly silent about one field.
Everything the crawler produces is a proposal
The pipeline writes into staging tables, never into the tables the app reads. A human reviews candidates and extractions through admin routes, including a duplicates view for rows that may be one programme twice. This keeps the model in the role it is good at — reading pages fast — and out of the role it is bad at, which is being trusted.
IVWhat it can’t do
Each of these is stated in the product itself, not only here.
Limitation
Coverage is only as wide as the seed list
Discovery finds programmes there is no row for, but it works from scored seeds and link classification — a funder whose site is JavaScript-rendered or paywalled simply will not appear. The absence of a scholarship from the dashboard says nothing about whether it exists.
Limitation
Grounding catches fabrication, not misreading
String-matching a quote proves the sentence is on the page. It does not prove the extractor understood it — a real quote can be attached to the wrong field. That is precisely why the human verification step exists and why "verified" is a status a crawler is forbidden to write.
Next
LinguaKu