Rendered at 07:09:13 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
palmotea 9 hours ago [-]
> The public posts and discussions being had about this subject are already informing AI companies on how to train their multimodal models to get around these obfuscations, most of which have already been broken. I'd argue every new font and tech demo is effectively a benchmark, daring AI firms come up with solutions to sidestep them. And they will be sidestepped, one way or another. If a human can see the information, that means there is a way the information can be parsed. "Ghost" fonts will become just another scraping obstacle with its own set of contingencies.
1. I don't like the sense of futility and powerlessness this advocates for.
2. I'm not sure it is so futile. I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.
That could happen if:
1. There are so many schemes out there the catalog of circumventions gets unwieldy.
2. Doubly so if the schemes allow generation of new obfuscated fonts per site or per page.
3. Then you're forcing the scrapers to pay a greater tax to get your text: spin up a Chrome instance to OCR a screenshot, or spend some a buck or two or LLM credits to reverse engineer the page in order to scrape it.
danudey 7 hours ago [-]
The main argument against these fonts is that you are removing accessibility for humans permanently in exchange for removing accessibility for scrapers temporarily.
It might raise the cost for some scrapers, to some degree, but it raises the cost to, effectively, infinity for anyone who needs to use a screen reader, a browser's 'reader' view, or any other assistive technology.
Things like this just remind me of EA's Spore; it released with DRM so draconian that legitimate players were getting locked out of the game during the first week, while people who pirated the game had no problem whatsoever and were playing the game without issue even before its release. The legitimate users of the thing were the only ones punished by the technology designed to stop everyone but them.
This is going to be the same thing; a website which AI will be able to read in short order but which assistive technologies will not be able to read ever.
jimmaswell 7 hours ago [-]
I also see no virtue in stopping an AI from reading something to begin with. If anything, I find it anti-social - hide your work from the machine trying to learn from it, taking nothing away from you in exchange for benefiting all of mankind?
r_lee 7 hours ago [-]
you mean benefit a private company training their AI model which they market to the public in order to get to their trillion dollar IPO?
satvikpendem 4 hours ago [-]
That's why one should advocate for open weight models.
infinite_spin 6 hours ago [-]
That seems like an unfair equivalence. Lots of things you enjoy, including this very forum, benefit already wealthy private companies. That doesn't seem like it's a very good litmus test for whether something is overall a benefit to humanity.
Retric 4 hours ago [-]
> overall a benefit to humanity.
By that metric a complete ban on LLM’s might be on the table, which I don’t think is something you’re advocating for here.
infinite_spin 2 hours ago [-]
I'm advocating for a litmus test that can't be easily abused, and I'm advocating against unfair equivalences. I also am not calling "overall a benefit to humanity" a metric, it's more of a conclusion we could arrive at by an application of relevant metrics.
Retric 1 hours ago [-]
> a conclusion we could arrive at by an application of relevant metrics.
That’s dangerous ground because of how you can arbitrarily change weights of different metrics. If the goal is the “overall benefit of humanity” then that’s what’s important not metrics.
infinite_spin 47 minutes ago [-]
How do you propose assessing "overall benefit of humanity" without metrics?
Retric 34 minutes ago [-]
Qualitative assessment etc not just metrics.
How do you propose to asses “overall benefit to humanity” with metrics?
infinite_spin 6 minutes ago [-]
Number of successful outcomes reported (e.g. with disabled students, cancer patients). Labor costs. Latency of services. Scientific discoveries found. Etc.
I'm not trying to discount the qualitative approach, but I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.
JodieBenitez 2 hours ago [-]
> taking nothing away from you
Hosting is not free.
silon42 1 hours ago [-]
That is the problem. I believe web is long overdue for a torrent like model where hosting is shared among all users (and ISPs instead of 'cloudflares').
brnt 39 minutes ago [-]
Opera Unite. The idea was that hosting HTML should be as simple as browsing it.
silon42 1 hours ago [-]
That is the problem. I believe web is long overdue for a torrent like model where hosting is shared among all users (and ISPs).
customguy 1 hours ago [-]
So scrapers are using those people as hostages. So let's have something like a HTML meta tag to link to an accessible version of a website and hard jail for using it for scraping.
It's not a technical issue, it's a social one, people behave differently given the same set of possibilities and incentives, and we can and should target those who fuck it up for everyone.
BrenBarn 3 hours ago [-]
I also don't like the fatalistic mindset, but I feel like we're better off trying to retaliate against the makers and operators of the bots, rather than getting into a technological arms race against the bots themselves. That is, the fight is one of policy, law, and morality, not of technology.
evnp 13 hours ago [-]
Thanks for introducing me to shieldfont.org! It's the first of these I've seen that feels designed to be more than a visual experiment, reading through their landing page is interesting. In particular, their section on accessibility seems to contradict this post's opening premise:
> Screen readers get the real words.
A screen reader reading down the page is never handed scrambled text, and our NVDA test asserts exactly that. Screen review and touch exploration are untested. ShieldFont hides shielded passages from accessibility tools by default with aria-hidden="true", because a decoy read aloud is fluent, grammatical, wrong English, and that is worse than silence.
> The real words remain sealed in the same page, and a visible notice above the block carries the control that uncovers them. It is on by default and reachable by mouse, keyboard and screen reader alike. Pressing it sets the reader’s browser to solving a compute-heavy puzzle: JavaScript and a few seconds of processing, more than most mass scrapers are willing to spend. That puts the words within reach of a screen reader, a translator and copy/paste.
I'd love to hear your thoughts on that. Also just aside, love this TUI-esque blog design and color palette (maybe a bearblog theme? still worth an upvote)
kstenerud 13 hours ago [-]
When you look at their live demo (https://shieldfont.org/demo/), it says: If you use a screen reader, custom font, or translator, please uncover the text before reading.
They also actively block copying the text, telling you to "uncover" the text first. The uncover operation is VERY expensive.
Anyone using assistive technologies or trying to copy "protected" text is SOL.
Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.
317070 9 hours ago [-]
> Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.
It's only search engines? And who uses those anymore anyway? Other bots?
It doesn't affect organic traffic, so you are not really putting yourself into a black hole. A lot of website are driven by social media and organic traffic, so they would be just fine with this approach.
8cvor6j844qw_d6 8 hours ago [-]
Same thoughts. The black hole concerns are overstated for personal stuff when nothing is at stake. Planning to adopt one of the fonts and see how it goes.
evnp 12 hours ago [-]
Isn't this the same sort of cost CloudFlare and anime catgirls are making us pay daily? Only enough to deter bots, or "a few seconds processing."
I take your point that one extra click/interaction required for screen readers only is objectively more friction, but it's only by exploring these technologies instead of dismissing them that we'll arrive at UX solutions truly work for people of all stripes (and ideally, not for bots, scrapers, and the like). I think if you're applying a tool like this you already have very different priorities than fueling Google.
kstenerud 12 hours ago [-]
As soon as you need to put a "decode" button for accessibility functions to work, you're effectively posting the key along with the cipher. It's self-defeating because any tool that supports screen readers will also support scraping. It's a fool's errand.
HeatrayEnjoyer 10 hours ago [-]
Accessibility isn't optional, so what do you propose?
kstenerud 6 hours ago [-]
I propose using the internet the way it was designed: Open and readable by everything (including bots). Scraping is a legal issue. There is no technical prevention mechanism that isn't theater.
svara 19 minutes ago [-]
A bit ironic that this is written in idiomatic Claudese.
csallen 12 hours ago [-]
It's funny, I was just thinking that the one thing I hate most about the terminal is that its monospaced fonts and overly-long lines are a nightmare for reading. So the fact that somebody designed a blog reading experience to mimic this is just… ugh for me.
But you seem to appreciate it, and I'm sure others do too. Different strokes for different folks.
condour75 14 hours ago [-]
Are these even meant to be used though? It seems more like performance art.
gruez 14 hours ago [-]
With this kinda of stuff it's hard to tell whether the person is doing it unironically, or knows it's "performative art". A while ago there was a trend of using a tool which imperceptibly perturbs an image in a way that supposedly breaks AI training on it. Of course, artists ate it up, despite the skepticism from AI researchers. Same with people setting up their sites to be "AI scraper traps", generating gibberish content. Probably also trivial to filter out, but people do it.
pixl97 11 hours ago [-]
The problem with being dumb satirically is dumb people look up to you as a thought leader.
yieldcrv 7 hours ago [-]
there was this guy that was anti-bitcoin - or specifically against most arguments from enthusiasts - and people started looking up to him to validate their feelings
and on news and podcasts he wound up correcting so many dumb arguments that he sounded pro-bitcoin and could never get to his own points
"well, no, not like that, the difficulty algorithm...."
"there are ways to use it with the power off"
"well, no, the transaction fees supplant the block reward so ..."
HlessClaudesman 46 minutes ago [-]
"Go ahead, obfuscate your contribution to the repository of all human knowledge, see if that impedes our imminent invasion! Moooahahahar!!!" - Kang and Kodos
ffsm8 10 hours ago [-]
"caveman speak" skill, need I say more?
People aren't particularly bright. That's why the scientific method was developed to counteract our built-in tendency for... Unorthodox approaches
Animats 9 hours ago [-]
Agreed. There was a thing a few years ago for dazzle-painting your face to avoid face recognition. This just makes you stand out.
blehn 13 hours ago [-]
The irony of championing accessibility using low-contrast simulated VGA text...
8 hours ago [-]
Narishma 13 hours ago [-]
Where do you see the low contrast?
smohare 13 hours ago [-]
The linked article has a fairly light gray background with white text. I can read it, but the low contrast is just tiresome.
VCFundedGenYer 13 hours ago [-]
Also the notion that the sides of the pages flash as you scroll due to simulating that old Macintosh monochrome monitor dithering effect. My eyes.
boxed 48 minutes ago [-]
Isn't that because you (and me!) have screens with slow response times for pixel color change though?
hn_throwaway_99 6 hours ago [-]
When I first was reading this I thought that the author was deliberately using a shitty font to make the point that obfuscated fonts are hard to read.
kazinator 3 hours ago [-]
I can tell right away that is stupid because one of the things I've used AI for was for deciphering someone's illegible handwriting, which it did amazingly.
Once it figures out for you what is written, you can't unsee it, so you know it has to be right.
Illegible fonts will only create accessibility problems for humans, even ones with normal vision, high literacy and no cognitive defects (dyslexia), while AI will blow right through the text.
If this is done in electronic documents, where the AI won't even see the glyps becaue it's reading the underlying character codes, it's even stupider.
I can't believe anyone would even try this (and then believe it is working without putting their hypotheses to the test).
frollogaston 4 hours ago [-]
Is the video on shieldfont.org AI-generated? It sounds like that voice again.
NetOpWibby 2 hours ago [-]
Who's gonna be the one to bring back Flash?
hk1337 14 hours ago [-]
Anti-AI fonts seems like scrambled porn on cable back in the 1980s.
piker 14 hours ago [-]
There could be benefits unlocked in legal documents by retaining a machine-readable version and distributing the obfuscated version with a legend at the top. We proposed one that said:
"This document contains mitigations against review by automated systems. Recipients should ensure that they have read the contents on screen or in print. Recipients with bona fide vision impairments may be entitled to unmitigated documents upon request."
In testing, obfuscating small portions of text slipped under the radar of most (then-)frontier LLMs.
We used a font that was rendered on the fly and reported faulty or fake Unicode mappings: https://tritium.legal/blog/noroboto but others have proposed and done the same with ligatures.
tbalsam 14 hours ago [-]
There was a story once about a boy with a wheelchair who needed a ramp to get into school, and the school made him use the loading dock ramp used for garbage and other things at the back. The school argued that it was an appropriate accommodation.
Accessibility is not accessible if you need to go through extra steps to get it.
Cerebrally, this as a solution makes sense. But if you know anyone with a vision or other impairment, gating it behind a request is not only cruel but gets within dangerous striking distance of an ADA lawsuit, for general applications.
Maybe in the legal field or specific niche cases it's possible. But this would represent a major step backwards in the work we've done lowering barriers for a population whose only difficulty in accessing common resources is because they were born, or got sick, differently than anyone else.
piker 14 hours ago [-]
My dad caught paralytic Polio at age 2 and has had limited mobility his entire life, so I'm familiar with that issue.
Our internal, hypothetical use-case was between contracting parties who were looking to avoid terms escaping into the wild. This shouldn't show up in standard ToS or similar. There are already really good legal reasons for that.
doctorpangloss 14 hours ago [-]
okay, i get that as a lawyer who wants to make money, every client is "heckin cute and valid." and you can hypothesize that this thing is something that clients want: "terms escaping into the wild," whatever that means - are you saying that you think copying and pasting an agreement into an LLM makes its contents escape into the wild, by some mechanism?
Look, I understand, you don't have to explain to me the theory for how that happens, I know it already. Since I know your a smart guy, to some extent you care about that only because you imagine that clients do. But in reality, in the real world, every email you send is read by at least two people, every contract you sign has multiple parties, etc. You make some obfuscated thing or whatever, but eventually, someone has the real text of the document - it might be YOUR client, it might be the person you are negotiating with, and you rarely represent ALL the sides. You never own ALL the information and all the parties and IT systems in totality. Eventually someone will put the text into a chatbot. Or maybe they put a salient piece of the pre-final text, like some legal theory or merely a question, into the chatbot.
So I see this font stuff, or watermarking stuff, or all this provenance and control stuff, as deeply illusory. It is the worst circlejerky kind of aesthetic experience making. When you mess with anti AI fonts you are trying to compete in the same business DocuSign is in, that is, in the business of selling holistic social experiences - a whole 7000 person company whose main competition is a fucking pen - but it's not like you're doing something creative. If you care about aesthetic experiences, write a short story! Are you getting it? The itch you are scratching with this weird thing, nobody wants.
stronglikedan 14 hours ago [-]
> Accessibility is not accessible if you need to go through extra steps to get it.
That seems a bit entitled to me, especially in the story you shared. They had access to the school just like everyone else. Why does it have to be in the exact same spot? Surely they could be dropped off by the loading dock as easily as others could be dropped off out front, and maybe even moreso. Should the wheelchair accessible stall be the first stall in the bathroom so they don't have to go through "extra steps" to get to it? As long as the ramp had the proper gradation to satisfy the ADA, I don't see the problem.
shnock 14 hours ago [-]
> They had access to the school just like everyone else
They literally did not. They had access to the school from a different entrance than everybody else. Specifically, one intended for cargo before people.
gblargg 13 hours ago [-]
They also literally cannot have the same access like everyone else, since they're in a wheelchair. Everything will be different.
fluoridation 14 hours ago [-]
That's a different sense of "just like". Not "in the same manner", but "to the same extent".
binaryturtle 14 hours ago [-]
Would it have been different, if they had everyone take the cargo entrance?
grim_io 13 hours ago [-]
Yes. It's about discrimination and dehumanizing everyday cruelty.
8 hours ago [-]
gizmo686 14 hours ago [-]
It is not sufficient to work against current AI. It needs to also work against AI that has been trained by a competent team aware of your mitigation. Or worse, a competent developer with no particular AI skills.
Otherwise, you are relying on obscurity, and will lose as soon as you become interesting enough to matter.
You will also break non-AI machine processing use cases. That isn't just accessibility, it is things like search.
piker 14 hours ago [-]
Yep
aleksejs 14 hours ago [-]
You will surely not have a good time enforcing the terms of a legal document that explicitly spells out that it is intentionally obfuscated from the party it intends to bind.
piker 14 hours ago [-]
No, that’s not at all what is going on here.
rappatic 6 hours ago [-]
Ironic that an article about bad fonts would use such an ugly, garish font
hartator 14 hours ago [-]
It also mostly don’t work.
bawolff 14 hours ago [-]
Yeah, it seems like most of these would probably be easier for an AI to read than a human once you give even a tiny bit of training to the AI.
there is a reason nobody uses text based captchas anymore.
dombiscoff 1 hours ago [-]
Quite frankly I don't understand why the cat and mouse argument here is meant to serve as a shutdown for anti-AI methods. Whole industries entirely exist in a cat and mouse state (cybersecurity, anti-cheat, etc) and no one in those industries imply that any solution is or can be a permanent dunk. If anything, the very fact that theres no permanent dunk is what leads to such industries developing a competitive service based industry to begin with. Why can't a theoretical anti-AI industry develop into the same thing? These fonts just seem like the infancy steps for such.
tabarnacle 6 hours ago [-]
Shieldfont approaches the accessibility issue mentioned by not obfuscating text on screen readers.
13 hours ago [-]
fluoridation 14 hours ago [-]
Wouldn't a font that shuffled the codepoint-to-glyph assignments be more effective?
CodesInChaos 13 hours ago [-]
That wouldn't make the decoy text look plausible to the AI.
Varelion 13 hours ago [-]
Is there evidence shieldfont doesn't work?
gs17 13 hours ago [-]
It only works while it's rare. If it was more common, scrapers would switch to OCR or simply reverse the font so they can decode the ligatures.
Their own whitepaper brings up a bigger issue: if it works, it poisons search engines as well.
Varelion 13 hours ago [-]
OCR is a more expensive option, right? I don't think these fonts need to stop ai from being trained; I think they just need to make it more expensive and difficult.
pixl97 10 hours ago [-]
More expensive and not worthwhile are two different things. Also when certain implementations become popular it's much more likely someone will write a very efficient kernel for decoding said text making it much less expensive than generic OCR.
Also more expensive doesn't mean that something won't happen, only the dynamics of how it happens. For example if you put all your documents in images then some service might just sell the AI providers the text. That service may do underhanded things like bundle OCR in an app that does something else and use your phone to get the text out of these images all day.
gs17 12 hours ago [-]
It's slightly more expensive, but the cost would be worth it if this became common. It's not designed to be hard to OCR, it's designed to be hard to copy out of the source code, so it doesn't require a very advanced OCR system (and all the regions that would need OCR-ing are clearly marked).
14 hours ago [-]
yieldcrv 7 hours ago [-]
I think this is an example of just catering to the gullible solely because the market exists without pondering anything about the individuals in the market
like the "pink tax", which isn't a tax at all but just a premium on consumer gullibility as the consumer can purchase other products that do the same thing simply marketed in a different way
playing into anti AI sentiment in a useless way fits the criteria
waffletower 15 hours ago [-]
I don't foresee anti-AI fonts being widely adopted. I see them largely as the symbolic saber rattling of intellectual property trolls.
dgellow 14 hours ago [-]
Actually, pretty sure it was an art performance
KPGv2 14 hours ago [-]
Every AI font proponent I've seen has been a fanfiction writer who just doesn't like AI stealing their shit to use against them.
dana-s 15 hours ago [-]
I believe the cat and rat game is already there, for multiple places, spam, captchas and now for AI content, yes, it's objective is to make it harder for AI companies to get such data, if it wastes their time, it's a win.
rpdillon 15 hours ago [-]
> it's objective is to make it harder for AI companies to get such data, if it wastes their time, it's a win.
That's only one half of the equation, though, isn't it? What if it makes it harder for legitimate users as well? It seems there's a balance to be struck.
pixl97 10 hours ago [-]
Kind of funny how you get downvoted for a rational take, but one that's not anti-ai.
You'd ask that same person how much they like captchas and I'm sure they'd think their a terrible idea and they've ran into all kinds of issues with them.
iwontberude 15 hours ago [-]
[dead]
aussieguy1234 5 hours ago [-]
It's probably trivial for AI to work around this either now or in the near future.
1. Screenshot page
2. Parse with OCR
Then train or do whatever else the font was trying to prevent...
anax32 9 hours ago [-]
Love that page styling.
ares623 6 hours ago [-]
Look at what they make us give.
grim_io 13 hours ago [-]
DRM for your eyes == garbage idea.
hellojomp 15 hours ago [-]
We are now in a weird middle ground where we want to write things OCR algorithms have trouble transcribing which also means we write things people with accessibility issues have trouble seeing. No child left behind?
mister_mort 15 hours ago [-]
It's like the old tale about the national park bin with the smartest bear / dumbest tourist crossover, except we're now comparing capabilities of the smartest AI with disabled humans.
Xirdus 14 hours ago [-]
It already was a major issue in 2010 - home desktop-grade OCRs could easily beat an average grandma on reading heavily garbled text.
no-name-here 14 hours ago [-]
> in 2010 - home desktop-grade OCRs could easily beat an average grandma
I’ve heard Tesseract OCR recommended, but even in 2026, no matter how I scan receipts or documents, the OCR output seems far worse than human reading?
strangecasts 13 hours ago [-]
I think the difficulty was specifically with CAPTCHA challenges, which had to be quick to generate but still legible - OCR on physical documents has to be robust to a different set of problems
That said, what kind of errors are you getting? I think the main practical difference between Tesseract's LSTM-based OCR and newer VLM-based OCR like PaddleOCR [1] is (hopefully) getting to skip making heuristics for the layout of the document, but errors possibly compounding over multiple tokens - are you getting individual illegible words or having the documents smooshed together because the OCR can't parse the layout?
Well, I used specialized software specifically for solving CAPTCHA on select few file sharing websites. My later experience with general purpose OCR software was just as awful as yours. But it's a product problem, not technology problem - you really don't need an advanced AI for it, you just need a good implementation of traditional shape-matching OCR that doesn't seem to exist anywhere for some reason.
avazhi 14 hours ago [-]
Not everything that annoys or inconveniences you is harmful, as if this needs to be said to an adult.
unethical_ban 15 hours ago [-]
Is it part of the joke that the site is intentionally over-pixelated while the author critiques readability? (edit: I don't mind esoteric design and I play old games. I found it funny to see a blog with aesthetics that are not optimized for long reading to complain about the readability of fonts)
I don't think any anti-AI font design is more than design-as-art statements against AI. If there is evidence of them being used exclusively for business and without accessible fonts aside them, I'm willing to be wrong.
I'm aware, and my comment isn't meant to be a challenge to anyone's sacred honor. I was observing a juxtaposition.
bradthebeaverfa 8 hours ago [-]
This author greatly overestimates how much I care about accessibility.
If I have to block a blind person from reading my blog to block an AI from training on it, I'll make that trade every time.
dombiscoff 18 minutes ago [-]
I think the line of reasoning that a method should provide as much accessibility as possible to not alienate real humans is a good one. You use 'have to' as if there aren't better solutions that could be invented, but such will only occur if we challenge the imperfect solutions we have now.
Ohentis 8 hours ago [-]
Well of the fonts listed, 2 out of the 3 would also prevent any human from wanting to read it and the third becomes an ineffective counter measure if it's widely used.
phoghed 4 hours ago [-]
I don’t understand how you think a font would block a bot from reading your blog in the first place?
It’s not even going to render the damn website. If it decides to and detects your retarded font it could just change the font trivially. You’d have to fundamentally fuck up the html text content for it to work at all, and then all you’ll do is inconvenience real people that you’d be extremely lucky to have attracted to your blog in the first place.
hyperadvanced 5 hours ago [-]
For real. It’s such a weak cop-out of an argument to lead with that I clicked out, never to read this blog again.
yipinwong 13 hours ago [-]
Not trynna to be funny.
The author's post is anti-human, thus useless and harmful (medically).
I can't read this bad font, sizing, spacing, etc.
The main offender is the color choice, and fonts that are just god aweful to read.
dombiscoff 17 minutes ago [-]
I think he just enjoys terminal styling, man.
dwrodri 13 hours ago [-]
Accessibility is important, and I find that "Reader Mode" in most browsers is quite good. Everyone should have access to the tools to consume content. Did your browser not provide that functionality?
yipinwong 12 hours ago [-]
basically make it look hard to read for the majority for the sake of 1%?
let those 1% use a diff tool view instead of 99% of us having to suffer.
That website no-way is accessible for my 30 year old eyes.
jotux 13 hours ago [-]
Found it awful to look at, tried to zoom out and everything on the page got larger. Seriously gross accessibility.
kokanee 13 hours ago [-]
I'm a bit frustrated by what seems to be a widespread strong negative reaction to anti-AI fonts. The accessibility problem is real, but I feel like that's a reason to push the investigation deeper for solutions to that problem, not a reason to abandon the effort entirely. The largest intellectual property infringement in the history of the universe is actively unfolding, and it's resulting in an existentially threatening transfer of wealth and power. That's a problem worth exploring every solution for, and solving it may entail some serious sacrifices.
Uberzi 13 hours ago [-]
It's simply that a font won't solve anything. That concept is worst than security by obscurity, as it causes more problem and add more constraints than what it solves... for a very limited time until AI bots are adjusted to decode those fonts properly.
waterTanuki 7 hours ago [-]
how can you make such a bold claim so early with 0 evidence? A successful (and much more realistic outcome) could be to make a font that is financially infeasible for LLMs to parse but easy for humans.
Every half assed means of trying to confuse an AI is just a small bit of learning away from making the AI better than you.
Worse when you have people with disabilities, which I seem to be this week, you just make doing things a pain in the ass.
What I don't get is people like you think there is a solution to this. There is not. The harder you try you either exclude more actual humans or you align the AI closer to how people actually see.
1. I don't like the sense of futility and powerlessness this advocates for.
2. I'm not sure it is so futile. I agree this stuff isn't encryption, which means it'll always be possible to circumvent the obfuscation, but it could raise the cost. Hopefully that can be done to the point where it's just not worth the bother.
That could happen if:
1. There are so many schemes out there the catalog of circumventions gets unwieldy.
2. Doubly so if the schemes allow generation of new obfuscated fonts per site or per page.
3. Then you're forcing the scrapers to pay a greater tax to get your text: spin up a Chrome instance to OCR a screenshot, or spend some a buck or two or LLM credits to reverse engineer the page in order to scrape it.
It might raise the cost for some scrapers, to some degree, but it raises the cost to, effectively, infinity for anyone who needs to use a screen reader, a browser's 'reader' view, or any other assistive technology.
Things like this just remind me of EA's Spore; it released with DRM so draconian that legitimate players were getting locked out of the game during the first week, while people who pirated the game had no problem whatsoever and were playing the game without issue even before its release. The legitimate users of the thing were the only ones punished by the technology designed to stop everyone but them.
This is going to be the same thing; a website which AI will be able to read in short order but which assistive technologies will not be able to read ever.
By that metric a complete ban on LLM’s might be on the table, which I don’t think is something you’re advocating for here.
That’s dangerous ground because of how you can arbitrarily change weights of different metrics. If the goal is the “overall benefit of humanity” then that’s what’s important not metrics.
How do you propose to asses “overall benefit to humanity” with metrics?
I'm not trying to discount the qualitative approach, but I think it's not impossibly hard to find metrics we can associate with "overall benefit to humanity" from a quantitative viewpoint.
Hosting is not free.
It's not a technical issue, it's a social one, people behave differently given the same set of possibilities and incentives, and we can and should target those who fuck it up for everyone.
> Screen readers get the real words. A screen reader reading down the page is never handed scrambled text, and our NVDA test asserts exactly that. Screen review and touch exploration are untested. ShieldFont hides shielded passages from accessibility tools by default with aria-hidden="true", because a decoy read aloud is fluent, grammatical, wrong English, and that is worse than silence.
> The real words remain sealed in the same page, and a visible notice above the block carries the control that uncovers them. It is on by default and reachable by mouse, keyboard and screen reader alike. Pressing it sets the reader’s browser to solving a compute-heavy puzzle: JavaScript and a few seconds of processing, more than most mass scrapers are willing to spend. That puts the words within reach of a screen reader, a translator and copy/paste.
I'd love to hear your thoughts on that. Also just aside, love this TUI-esque blog design and color palette (maybe a bearblog theme? still worth an upvote)
They also actively block copying the text, telling you to "uncover" the text first. The uncover operation is VERY expensive.
Anyone using assistive technologies or trying to copy "protected" text is SOL.
Search engines will index the decoy. You'll get no traffic. Their solution is basically putting yourself in a black hole.
It's only search engines? And who uses those anymore anyway? Other bots?
It doesn't affect organic traffic, so you are not really putting yourself into a black hole. A lot of website are driven by social media and organic traffic, so they would be just fine with this approach.
I take your point that one extra click/interaction required for screen readers only is objectively more friction, but it's only by exploring these technologies instead of dismissing them that we'll arrive at UX solutions truly work for people of all stripes (and ideally, not for bots, scrapers, and the like). I think if you're applying a tool like this you already have very different priorities than fueling Google.
But you seem to appreciate it, and I'm sure others do too. Different strokes for different folks.
and on news and podcasts he wound up correcting so many dumb arguments that he sounded pro-bitcoin and could never get to his own points
"well, no, not like that, the difficulty algorithm...."
"there are ways to use it with the power off"
"well, no, the transaction fees supplant the block reward so ..."
People aren't particularly bright. That's why the scientific method was developed to counteract our built-in tendency for... Unorthodox approaches
Illegible fonts will only create accessibility problems for humans, even ones with normal vision, high literacy and no cognitive defects (dyslexia), while AI will blow right through the text.
If this is done in electronic documents, where the AI won't even see the glyps becaue it's reading the underlying character codes, it's even stupider.
I can't believe anyone would even try this (and then believe it is working without putting their hypotheses to the test).
"This document contains mitigations against review by automated systems. Recipients should ensure that they have read the contents on screen or in print. Recipients with bona fide vision impairments may be entitled to unmitigated documents upon request."
In testing, obfuscating small portions of text slipped under the radar of most (then-)frontier LLMs.
We used a font that was rendered on the fly and reported faulty or fake Unicode mappings: https://tritium.legal/blog/noroboto but others have proposed and done the same with ligatures.
Accessibility is not accessible if you need to go through extra steps to get it.
Cerebrally, this as a solution makes sense. But if you know anyone with a vision or other impairment, gating it behind a request is not only cruel but gets within dangerous striking distance of an ADA lawsuit, for general applications.
Maybe in the legal field or specific niche cases it's possible. But this would represent a major step backwards in the work we've done lowering barriers for a population whose only difficulty in accessing common resources is because they were born, or got sick, differently than anyone else.
Our internal, hypothetical use-case was between contracting parties who were looking to avoid terms escaping into the wild. This shouldn't show up in standard ToS or similar. There are already really good legal reasons for that.
Look, I understand, you don't have to explain to me the theory for how that happens, I know it already. Since I know your a smart guy, to some extent you care about that only because you imagine that clients do. But in reality, in the real world, every email you send is read by at least two people, every contract you sign has multiple parties, etc. You make some obfuscated thing or whatever, but eventually, someone has the real text of the document - it might be YOUR client, it might be the person you are negotiating with, and you rarely represent ALL the sides. You never own ALL the information and all the parties and IT systems in totality. Eventually someone will put the text into a chatbot. Or maybe they put a salient piece of the pre-final text, like some legal theory or merely a question, into the chatbot.
So I see this font stuff, or watermarking stuff, or all this provenance and control stuff, as deeply illusory. It is the worst circlejerky kind of aesthetic experience making. When you mess with anti AI fonts you are trying to compete in the same business DocuSign is in, that is, in the business of selling holistic social experiences - a whole 7000 person company whose main competition is a fucking pen - but it's not like you're doing something creative. If you care about aesthetic experiences, write a short story! Are you getting it? The itch you are scratching with this weird thing, nobody wants.
That seems a bit entitled to me, especially in the story you shared. They had access to the school just like everyone else. Why does it have to be in the exact same spot? Surely they could be dropped off by the loading dock as easily as others could be dropped off out front, and maybe even moreso. Should the wheelchair accessible stall be the first stall in the bathroom so they don't have to go through "extra steps" to get to it? As long as the ramp had the proper gradation to satisfy the ADA, I don't see the problem.
They literally did not. They had access to the school from a different entrance than everybody else. Specifically, one intended for cargo before people.
Otherwise, you are relying on obscurity, and will lose as soon as you become interesting enough to matter.
You will also break non-AI machine processing use cases. That isn't just accessibility, it is things like search.
there is a reason nobody uses text based captchas anymore.
Their own whitepaper brings up a bigger issue: if it works, it poisons search engines as well.
Also more expensive doesn't mean that something won't happen, only the dynamics of how it happens. For example if you put all your documents in images then some service might just sell the AI providers the text. That service may do underhanded things like bundle OCR in an app that does something else and use your phone to get the text out of these images all day.
like the "pink tax", which isn't a tax at all but just a premium on consumer gullibility as the consumer can purchase other products that do the same thing simply marketed in a different way
playing into anti AI sentiment in a useless way fits the criteria
That's only one half of the equation, though, isn't it? What if it makes it harder for legitimate users as well? It seems there's a balance to be struck.
You'd ask that same person how much they like captchas and I'm sure they'd think their a terrible idea and they've ran into all kinds of issues with them.
1. Screenshot page
2. Parse with OCR
Then train or do whatever else the font was trying to prevent...
I’ve heard Tesseract OCR recommended, but even in 2026, no matter how I scan receipts or documents, the OCR output seems far worse than human reading?
That said, what kind of errors are you getting? I think the main practical difference between Tesseract's LSTM-based OCR and newer VLM-based OCR like PaddleOCR [1] is (hopefully) getting to skip making heuristics for the layout of the document, but errors possibly compounding over multiple tokens - are you getting individual illegible words or having the documents smooshed together because the OCR can't parse the layout?
[1] https://github.com/PaddlePaddle/PaddleOCR
I don't think any anti-AI font design is more than design-as-art statements against AI. If there is evidence of them being used exclusively for business and without accessible fonts aside them, I'm willing to be wrong.
If I have to block a blind person from reading my blog to block an AI from training on it, I'll make that trade every time.
It’s not even going to render the damn website. If it decides to and detects your retarded font it could just change the font trivially. You’d have to fundamentally fuck up the html text content for it to work at all, and then all you’ll do is inconvenience real people that you’d be extremely lucky to have attracted to your blog in the first place.
I can't read this bad font, sizing, spacing, etc. The main offender is the color choice, and fonts that are just god aweful to read.
That website no-way is accessible for my 30 year old eyes.
Every half assed means of trying to confuse an AI is just a small bit of learning away from making the AI better than you.
Worse when you have people with disabilities, which I seem to be this week, you just make doing things a pain in the ass.
What I don't get is people like you think there is a solution to this. There is not. The harder you try you either exclude more actual humans or you align the AI closer to how people actually see.