# Examples of curlminder usage

**URL:** <https://forum.beeminder.com/t/examples-of-curlminder-usage/12040>\
**Category:** Akrasia\
**Created:** [December 8, 2024, 4:20pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040 "2024-12-08T16:20:37Z")\
**Posts on this page:** 16\
**Page:** 1

<div class="post-metadata">

**Author:** ![clivemeister](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/clivemeister/32/2034_2.png) [@clivemeister](https://forum.beeminder.com/u/clivemeister)\
**Post date:** [December 8, 2024, 4:20pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/1 "2024-12-08T16:20:37Z")

</div>

We’ve just launched the curlminder integration, so I thought it might be useful to gather a few examples for people. Please contribute your own - or maybe ask questions about one, if you’d like to try it out but don’t speak regex well enough!

Here’s an example of one that doesn’t work right now, but - I hope - we can get to work with our joint efforts! I want to extract my XP points score from Chessable. Here’s the URL: [Chessable](https://www.chessable.com/profile/clive_f) And here’s how the relevant segment looks in my browser:  
 ![image](https://us1.discourse-cdn.com/flex019/uploads/beeminder/original/2X/7/73d5720edc905818fdc436a775e1cc5b249533d7.png)

Now I can read the html source in my browser, and could probably cook up the necessary regex… but right now, when I start setting up the goal, when it shows me the initial html in Beeminder it ends with “You have been blocked If you are using a VPN, disconnect and try again”. So once this has been fixed (if it can be fixed!), I’ll have another go.

But meantime, do post your own url+regex recipes here, for others to try!

---

<div class="post-metadata">

**Author:** ![baronvonchickenpants](https://avatars.discourse-cdn.com/v4/letter/b/d26b3c/32.png) [@baronvonchickenpants](https://forum.beeminder.com/u/baronvonchickenpants)\
**Post date:** [December 8, 2024, 5:45pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/2 "2024-12-08T17:45:38Z")

</div>

Thanks for starting this thread off @clivemeister

---

<div class="post-metadata">

**Author:** ![shanaqui](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/shanaqui/32/3513_2.png) [@shanaqui](https://forum.beeminder.com/u/shanaqui)\
**Post date:** [December 9, 2024, 1:53pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/3 "2024-12-09T13:53:36Z")

</div>

Mine was written for me by @dreev so I could experiment! It pulls my number of minions collected in the game Final Fantasy XIV from [this page](https://na.finalfantasyxiv.com/lodestone/character/34992495/minion/), using this expression:

`Total:\s*\<span\>\s*([\d\,]+)`

Other stuff I might be tempted to do in the future, all FFXIV-related:

- Number of mounts: [from here](https://na.finalfantasyxiv.com/lodestone/character/34992495) [can probably use the same regex on this URL, I think?]
- Number of achievements: [from this page](https://na.finalfantasyxiv.com/lodestone/character/34992495/achievement/)
- Achievement points: ditto
- Levequest completion: [via FFXIVcollect](https://ffxivcollect.com/characters/34992495)
- Relic completion: ditto

Buuut that’d require help or me learning what all this stuff is about in a lot more detail, so it’s not a current project. 😛

---

<div class="post-metadata">

**Author:** ![felixm](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/felixm/32/4955_2.png) [@felixm](https://forum.beeminder.com/u/felixm)\
**Post date:** [December 9, 2024, 10:59pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/4 "2024-12-09T22:59:31Z")

</div>

It could be helpful to specify which regex flavor/style Curlex supports in the documentation ([Curlex - Beeminder Help](https://help.beeminder.com/article/364-curlex)). It’s likely Perl-style, but making this explicit might help users, especially with GPT requests.

Speaking about GPT, the process to create your own regexes might be relatively straightforward, @shanaqui.

1. Go to the site you care about, I used your first example ([Eirian Evanna | FINAL FANTASY XIV, The Lodestone](https://na.finalfantasyxiv.com/lodestone/character/34992495/mount/)), but clicked on “Mounts,” so this is the URL I used: [Eirian Evanna | FINAL FANTASY XIV, The Lodestone](https://na.finalfantasyxiv.com/lodestone/character/34992495/mount/).
2. Right click and _Inspect_ the number you care about. It should then look something like this:
3. ![image](https://us1.discourse-cdn.com/flex019/uploads/beeminder/original/2X/d/dd583cb694385bde84a669a6f7a42d2b8916000c.png)
4. Now, right click on an outer tag that includes your target number and select _copy outer HTML_. You have to apply your best judgement to know which tag to use. `<span>` is probably not enough context, but `class="minion __sort__ total"` kind of looks like we have enough context to unambiguously identify the position.
5. It might then look like this when you paste it:
6. `<p class="minion __sort__ total">Total: <span>201</span></p>`
7. Now, you can ask GPT to create the regex for you:

Please create a Perl-style regex that matches the following HTML code. Please include a match group that matches the number that is part of the HTML. The number might change and there should be only one match group for that number.

```auto
<p class="minion __sort__ total">Total: <span>201</span></p>

```

Claude 3.5 Sonnet then gives me this:

`<p class="minion __sort__ total">Total: <span>(\d+)</span></p>`

And GPT 4o this:

`/<p class="minion __sort__ total">Total: <span>(\d+)<\/span><\/p>/`

Here, you will have to remove the slashes before pasting it into Beeminder.

As a sanity check, I will try your second example: [Eirian Evanna | FINAL FANTASY XIV, The Lodestone](https://na.finalfantasyxiv.com/lodestone/character/34992495/achievement/).

Right click, inspect:

![image](https://us1.discourse-cdn.com/flex019/uploads/beeminder/original/2X/6/6e326368c8deb13fcea1956a69adb40be814bde3.png)

Select marked HTML tag for enough context, and copy outer:

```auto
						<div class="select-pulldown en-us">
								<form action="?">
									<select name="order" onchange="this.form.submit()" class="select-pulldown__open">
										<option value="1">Sort by most recent</option>
										<option value="2">Sort by oldest</option>
									</select>
								</form>
						</div>
						<div class="parts__total">2117 Total</div>
					</div>

```

Paste into Claude with prompt; I guess it didn’t agree that we need that much context:

```auto
<div class="parts__total">(\d+) Total</div>

```

Test in Beeminder:

 ![image](https://us1.discourse-cdn.com/flex019/uploads/beeminder/original/2X/1/1061636e35a71b6a18a71181cbafe2a8dafc050a.png)

Edit:

For achievement points: `<p class="achievement__point">(\d+)</p>`

Also: obligatory disclaimer. I don’t endorse “parsing” HTML with regexes and this shouldn’t be used for anything serious. (See first answer here:  
[html - RegEx match open tags except XHTML self-contained tags - Stack Overflow](https://stackoverflow.com/questions/1732348/regex-match-open-tags-except-xhtml-self-contained-tags))

---

<div class="post-metadata">

**Author:** ![dreev](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/dreev/32/4311_2.png) [@dreev](https://forum.beeminder.com/u/dreev)\
**Post date:** [December 12, 2024, 12:32am UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/5 "2024-12-12T00:32:32Z")

</div>

> [@clivemeister](#):
>
> Now I can read the html source in my browser, and could probably cook up the necessary regex… but right now, when I start setting up the goal, when it shows me the initial html in Beeminder it ends with “You have been blocked If you are using a VPN, disconnect and try again”. So once this has been fixed (if it can be fixed!), I’ll have another go.

I have failed to replicate that specifically. When I do view-source on the page and grep for your XP it seems to simply be absent. So I think this is fundamentally the same problem described [in the blog post](https://blog.beeminder.com/curlminder) for Fatebook, where the number is populated by Javascript. (Getting blocked from fetching the html would then be an orthogonal problem…)

PS: Huge thanks to @felixm for the tutorial on getting LLMs to help with the regex part. Now I’m getting tempted to make a muggle-friendly version of Curlminder where you can just describe the number on the page you want to beemind in English and Beeminder makes the calls to the LLM to construct the regex. But I think we need to collect more use cases before we can justify that. (Keep em coming, y’all!)

---

<div class="post-metadata">

**Author:** ![clivemeister](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/clivemeister/32/2034_2.png) [@clivemeister](https://forum.beeminder.com/u/clivemeister)\
**Post date:** [December 12, 2024, 9:35am UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/6 "2024-12-12T09:35:41Z")

</div>

> [@dreev](#):
>
> I have failed to replicate that specifically. When I do view-source on the page and grep for your XP it seems to simply be absent.

I’ve tried using curl from the command line, and I get back a page with stuff saying things like “This website is using a security service to protect itself from online attack… you can email the site owner to let them know you were blocked”. So I did!

I suspect if I fiddled about enough with curl settings, to make myself look more like a browser, I could get back the page. Then we could see if it’s a Javascript-populated thing or not. I may have a go at some point… or just ask Gemini for suggestions, probably!

---

<div class="post-metadata">

**Author:** ![baronvonchickenpants](https://avatars.discourse-cdn.com/v4/letter/b/d26b3c/32.png) [@baronvonchickenpants](https://forum.beeminder.com/u/baronvonchickenpants)\
**Post date:** [December 12, 2024, 9:46am UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/7 "2024-12-12T09:46:30Z")

</div>

I think this is going to be a bit of both problems

The message you get when trying to set up a curl request looks to me like that request is being blocked by cloudfire at chessable’s end. However, even when accessing it via a full browser the XP figure is not available when viewing the page source so is most likely getting populated by a script.

If that’s the case then I dont think solving the ‘blocking’ issue will make the XP figure available to curl

---

<div class="post-metadata">

**Author:** ![aad](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/aad/32/4736_2.png) [@aad](https://forum.beeminder.com/u/aad)\
**Post date:** [December 21, 2024, 8:43am UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/8 "2024-12-21T08:43:37Z")

</div>

I set up a goal to track my [Criticker.com](http://Criticker.com) ratings as a proxy for how many movies I watch. With my cinema subscription restarted, this helps me make sure I’m not paying too much for the subscription.

Criticker profiles show the review count in plain text:

1. Use your public profile URL: `https://www.criticker.com/profile/<username>`
2. Extract the count with: `>(\d+)\sFilm\sRatings<`

---

<div class="post-metadata">

**Author:** ![dreev](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/dreev/32/4311_2.png) [@dreev](https://forum.beeminder.com/u/dreev)\
**Post date:** [December 22, 2024, 10:31pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/9 "2024-12-22T22:31:30Z")

</div>

Nice! I hadn’t heard of Criticker but know someone who uses Letterboxd, which seems similar and also works with Curlminder!

> **[Zvi Mowshowitz’s films](https://letterboxd.com/thezvi/films/)**
>
> Zvi Mowshowitz’s films

The number of watched and reviewed films are in tooltips but Curlminder can handle that fine.

For example, you can find a chunk of html for number of watched films that looks like this:

`title="102&nbsp;films">Watched</a>`

Which you can turn into a regex like so:

`title\=\"([\d\,]+)[^\d]+?(?i)films.+watched`

---

<div class="post-metadata">

**Author:** ![baronvonchickenpants](https://avatars.discourse-cdn.com/v4/letter/b/d26b3c/32.png) [@baronvonchickenpants](https://forum.beeminder.com/u/baronvonchickenpants)\
**Post date:** [January 31, 2025, 2:33pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/10 "2025-01-31T14:33:42Z")

</div>

Did we ever figure out a way to scrape the Chessable XP site?

---

<div class="post-metadata">

**Author:** ![clivemeister](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/clivemeister/32/2034_2.png) [@clivemeister](https://forum.beeminder.com/u/clivemeister)\
**Post date:** [February 9, 2025, 1:37pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/11 "2025-02-09T13:37:50Z")

</div>

I did a bit of experimenting, and even when using curl with a bunch of plausible headers, it looks like it won’t work unless you’re doing the GET from a real browser. There are various ways to do this on the server side (e.g. Capybara), but they’re all relatively expensive (in cpu cycles/memory), and not super-reliable, so I don’t know if they’re terribly practical.

---

<div class="post-metadata">

**Author:** ![philip](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/philip/32/1505_2.png) [@philip](https://forum.beeminder.com/u/philip)\
**Post date:** [February 11, 2025, 8:52pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/12 "2025-02-11T20:52:57Z")

</div>

Interesting! Thanks to this thread I’ve tried both Criticker and now, today, Letterboxd. The latter has an import feature which sounded tempting but the data quality seems poor. So far 32% of my watched titles don’t exist there, even when they’re present in the underlying tmdb database – and those tmdb ID’s are wrongly mapped to other titles.

**Update:** I guess Letterboxd is obsessed with whatever they think of as films, because most of the missing titles (58/192) are series of one kind or another. But not all series in my list are missing, and not all movies are present. The worst thing from my pov is that if you feed it an ID from their official data source, tmdb, it returns a confident (but wrong!) result rather than a not-found.

**Update the second:** turns out that tmdb distinguishes between tv and film, using different ID sequences, so Letterboxd was indeed finding the film numbered X when I meant the series. Confusingly, some TV series are present on Letterboxd. The response from support was that they “don’t currently support cancelled TV series on the platform”. So, despite the slicker presentation of Letterboxd, I think I’ll be sticking with Criticker.

---

<div class="post-metadata">

**Author:** ![aad](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/aad/32/4736_2.png) [@aad](https://forum.beeminder.com/u/aad)\
**Post date:** [March 24, 2025, 3:12pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/13 "2025-03-24T15:12:58Z")

</div>

Criticker changed design. This seems to work now:

This URL is easier now: [Film Database - Recommendations & Reviews | Criticker](https://www.criticker.com/ratings/)/  
With this expression: `of\s+(\d+)\s+Titles`

---

<div class="post-metadata">

**Author:** ![philip](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/philip/32/1505_2.png) [@philip](https://forum.beeminder.com/u/philip)\
**Post date:** [March 28, 2025, 11:33am UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/14 "2025-03-28T11:33:27Z")

</div>

I’ve just hooked up criticker to curlminder, thank you!

> [@aad](#):
>
> With this expression: `of\s+(\d+)\s+Titles`

A few things to add:

The public URL includes your criticker username: `https://criticker.com/ratings/<username>`

Go into the goal’s settings tab and untick the `cumulative` box near the bottom of the page; this is an odometer-style goal, where the current total is posted as the datapoint.

If you want to import some historical data, you can email [bot@beeminder.com](mailto:bot@beeminder.com) with a subject of `<username>/<goalname>` and datapoints in the body formatted as `yyyy mm dd value`.

I asked my friendly neighbourhood LLM to summarise the recent ratings copied from my public profile page in a short conversation, something like:

- how many entries are on each date?
- could you format those as yyyy mm dd value ?
- cumulative values please
- no, cumulative the other way!
- recalculate the cumulative values so that the most recent total is X

So much faster that either counting them or scripting it.

---

<div class="post-metadata">

**Author:** ![divide](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/divide/32/6162_2.png) [@divide](https://forum.beeminder.com/u/divide)\
**Post date:** [June 10, 2025, 1:06pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/15 "2025-06-10T13:06:37Z")

</div>

Total DuoLingo XP:  
URL: [https://www.duolingo.com/2017-06-30/users?username=](https://www.duolingo.com/2017-06-30/users?username=)  
regex: `/"totalXp":(\d+)/`

---

<div class="post-metadata">

**Author:** ![philip](https://sea1.discourse-cdn.com/flex019/user_avatar/forum.beeminder.com/philip/32/1505_2.png) [@philip](https://forum.beeminder.com/u/philip)\
**Post date:** [October 6, 2026, 5:00pm UTC](https://forum.beeminder.com/t/examples-of-curlminder-usage/12040/16 "2026-10-06T17:00:51Z")

</div>

> [@aad](#):
>
> Criticker changed design.

Again. They’ve gone hard down the route of anti-scraping on account of the aggressive bots.

But it seems that we were doing it wrong. Criticker provides a feed, so RSSMinder is your friend.

Many feeds, in fact: [RSS Feeds And Film Rating Import And Exports | Criticker](https://www.criticker.com/resources/)
