BrowserBook | Not Hacker News!

Discussion (30 comments)

Showing 37 comments

22 days ago

1 reply

Interesting. Quick question in regards to the code generation : Do you dump the DOM to provide relevant context to build the automation or does the agent automatically tries to discover relevant segments (like a claude code) ?

cschlaepfer

22 days ago

Appreciate it!

Yes exactly - today we send it a simplified version of the DOM, but we're currently building an agent which will be able to discover the relevant DOM elements.

huntaub

22 days ago

2 replies

This is a super interesting product, guys. I get that agents aren't great for everything right now, but I'd expect that they'll continue to improve over time (like everything in the LLM space).

How do you see the product evolving as agents become better and better?

cschlaepfer

22 days ago

1 reply

Thanks, and great question - we think about this a lot and think there are a couple of things here.

First, as models get better, our agent's ability to navigate a website and generate accurate automation scripts will improve, giving us the ability to more confidently perform multi-step generations and get better at one-shotting automations.

We expect browser agents will improve as well, which I think is more along the lines of what you're asking. At scale, we still think scripts will be better for their cost, performance, and debuggability aspects - but there are places where we think browser agents could potentially fit as an add-on to deterministic workflows (e.g., handling inconsistent elements like pop-ups or modals). That said, if we do end up introducing a browser agent in the execution runtime, we want to be very opinionated about how it can be used, since our product is primarily focused on deterministic scripting.

huntaub

22 days ago

> At scale, we still think scripts will be better for their cost, performance, and debuggability aspects

This actually makes a ton of sense to me in lots of the LLM contexts (e.g. seeing how we are starting to prefer having LLMs write one-off scripts to do API calls rather than just pointing them at problems and having them try it directly).

Thanks!

smt88

22 days ago

Agents have plateaued and will never be as good as deterministic code for browser automation. It's a fundamental issue with the way LLMs work.

timerol

22 days ago

2 replies

> only available for Mac so far, sorry!

Is there a plan to change this? Building on Electron should make it manageable to go cross-platform. 2026 will be the year of the Linux desktop, as the prophecies have long foretold.

Off-topic, but Kernel refers to https://www.onkernel.com. A bit of an awkward name

koakuma-chan

22 days ago

1 reply

Why would something be only available for Mac? What are they using?

dang

22 days ago

1 reply

In the early stage of a product the most important thing is fast iteration loops, so it can be wise to postpone later projects (such as porting to other platforms) - things which are important, but can come later - in order to focus on what the product should be.

koakuma-chan

22 days ago

I am just wondering from technical standpoint. Their product doesn't seem like it would be doing anything OS-specific, so I would expect that it would be multi-platform without them having to do anything.

jorrie

22 days ago

Yes! Going cross-platform is on the roadmap, but right now we're focused on building out the feature set and improving the form factor.

Also yeah, Kernel refers to onkernel - they call it that because they're running the browsers in a unikernel. They're a great product and you should give them a look if you need hosted browsers!

poly2it

22 days ago

2 replies

I feel like I've seen a product similar to this quite recently on HN, but it was a standalone agentic workflow which was open source on GitHub. I can't seem to find it right now. Does anyone know what I'm referring to?

jasonjmcghee

22 days ago

There have been a bunch

https://hn.algolia.com/?dateRange=pastYear&page=0&prefix=fal...

suchintan

22 days ago

maybe https://github.com/Skyvern-AI/Skyvern ?

jackconsidine

22 days ago

2 replies

Congrats on the launch

Please port to linux soon (sure it's relatively trivial on Electron :)).

Like the idea of the IDE. Seems like it'd make it easy to prototype and launch quickly.

RE: embrace the suck, yeah I'm with you. I prefer the brittleness of scripts to non-deterministic (potentially unhinged) workflows

cschlaepfer

22 days ago

1 reply

Thanks! Yep, linux is coming soon - now that we have the first version of the IDE out the door we're going to get cross-platform going shortly.

dman

22 days ago

1 reply

What UI framework are you using to build the app?

cschlaepfer

22 days ago

2 replies

It's a typescript/react application, bundled in electron.

dman

22 days ago

thank you!

rtolsma

22 days ago

is windows support going to be included too since it's electron?

kyriakos

22 days ago

same.. I'm on Windows.. clicked download only to realize its mac only.

devmor

22 days ago

1 reply

Wow, really cool project. As someone who's not primarily a frontend developer but has had to write a lot of browser-based feature tests, I love the concept and execution.

Why the subscription model though? That's the one thing that concerns me.

Is data being sent back to your servers to enable some of the functionality? I don't speak for my employer here, but as someone who works in the healthcare technology industry, if I wanted to get my bosses to buy into this, I would be looking for something that we license to run on our environments.

cschlaepfer

22 days ago

Thank you!

> Why the subscription model though?

The subscription model is primarily to cover the costs of creating and running automations at scale (i.e., LLM code gen and browser uptime) and to build a sustainable business around those features. We included the free tier to give users access to the IDE, but we're committed to adding value beyond just the IDE and subscriptions support that.

> Is data being sent back to your servers to enable some of the functionality?

Yes - we save all notebooks in our database. Since we're working to build a lot of value-add features for hosted executions, having notebooks saved online worked in service of that.

That said, we're now thinking about the local-only / no sign-up use case as well. We've gotten a lot of feedback about this, so it's something we're taking seriously now that we've gotten all of the core functionality in.

I will also add, we are HIPAA-compliant for healthcare use cases.

Really appreciate the questions!

jackienotchan

22 days ago

1 reply

Congrats! Could this also be used to generate e2e test automations? For scraping, how do you handle Cloudflare and Captchas? Do you respect robots.txt instructions of websites?

cschlaepfer

22 days ago

Thanks, we appreciate it!

Yes, you can use BrowserBook to write e2e test automations, but we don't currently include playwright assertions in the runtime - we excluded these since they are geared toward a specific use case, and we wanted to build more generally. Let us know if you think we should include this though; we're always looking for feedback.

> For scraping, how do you handle Cloudflare and Captchas?

Cloudflare turnstiles/captchas tend to be less of an issue in the inline browser because it’s just a local Chrome instance, avoiding the usual bot-detection flags from headless or cloud browsers (datacenter IPs, user-agent quirks, etc.). For hosted browsers, we use Kernel's stealth mode to similar effect.

> Do you respect robots.txt instructions of websites?

We leave this up to the developer creating the automations.

orliesaurus

22 days ago

1 reply

I've been hacking together my own browser automations... the idea of deterministic scripting resonates with me... but I'm wondering how BrowserBook plans to handle authentication flows that require 2FA or CAPTCHAs.

ALSO is there any plan for integrating with CI pipelines... being able to run these scripts headless on servers would be huge.

BUT overall it's refreshing to see someone lean into brittle scripts rather than hide behind agent magic...

cschlaepfer

22 days ago

BrowserBook allows users to create 'auth profiles' which can be utilized in notebooks for authentication purposes. These profiles currently support username/password and 2FA via TOTP (and we recommend provisioning a service account for your automations).

For captchas, we use Kernel's stealth mode which includes a captcha solver.

Re: CI integration, today we support API-based execution, but if you have a specific CI pipeline or set of tools you'd like to see support for, let us know and we can look into it!

nickstaggs

22 days ago

1 reply

Awesome product! I really liked how the auth profiles work as well. While the primary use case is workflow automation are there any roadmap items on integrating this with the developer experience? A previous company I was at was fairly fond of e2e tests in playwright and this seems like it would have been a huge boon for writing them quickly.

cschlaepfer

22 days ago

Thanks!

We're constantly thinking about ways we can improve the dev experience and integrations story around deploying these scripts. Right now we support API executions, and we are adding webhooks soon. We think this will unblock the earliest adopters, and as we learn more about popular use cases/workflows, we'll look to prioritize first-class integrations where it makes sense.

innagadadavida

22 days ago

1 reply

I'm curious why use a hosted browser instead of just spinning one up locally and since you already have he electron app. Why not just use a different Chrome profile for isolation and interact with that?

cschlaepfer

22 days ago

Thanks for the question! We only use the hosted browser for running the automations remotely (via API). In the IDE, we use a local chrome browser, where we spin up an anonymous profile for isolation.

rcarmo

22 days ago

I like this and find it profoundly weird in equal parts. I get the use case and why it’s being done in the browser, but like RPA and similar tech, I have to wonder at the path that the industry took to get here and make this a viable and clever solution.

Somewhere there is a timeline where front-ends evolved differently.

danecjensen

21 days ago

How long until cursor copies this

asdev

22 days ago

Are you still doing healthcare back office automation? Would love to learn why you pivoted out if not, happy to DM as well

Imustaskforhelp

20 days ago

Is there a way to move code outside of browserbook, I tried it and built something and wanted to deploy it on my vps or similar so I wished to move the code out but I am sorry but I tried pupeeteer, playwright etc. to work with it but it simply doesnt work in the context of pupeteer,playwright similar and I would prefer if there was a way to just run the code itself/move the code perhaps to something which can run on vps or similar?

radial_symmetry

22 days ago

Cool product, but this post is too long and it's hard to find your link in it.

witnessme

22 days ago

This is a novel idea. Somewhere between the extremes of being useful vs being an overkill. More towards overkill because of its dependency on a new app/browser that needs to be installed. But I'm looking forward to more development on this idea, making it a production-ready automation.

Resources