For a school activity we are organizing a rubber band bracelet making. Here is a how to we did for it. Feel free to reuse it if you’re doing the same.

For a school activity we are organizing a rubber band bracelet making. Here is a how to we did for it. Feel free to reuse it if you’re doing the same.

LLMs have seen a huge surge of popularity with ChatGPT by going from prompt to text for various use cases. But what’s really exciting is that they are also extremely useful as ways to implement normal functions within a program. This is what I call LLM As A Function.
For example, if you want to build a website builder that will only use components that you have in your design system, you can write the following:
const prompt = 'Build a chat app UI'; const components = llm<Array<string>>( 'You only have the following components: ' + designSystem.getAllExistingComponents().join(', ') + 'n' + 'What components to do you need to do the following:n' + prompt ); // ['List', 'Card', 'ProfilePicture', 'TextInput'] const result = llm<{javascript: string, css: string}>( 'You only have the following components: ' + components.join(',') + 'n' + 'Here are examples of how to use them:n' + components.map(component => designSystem.getExamplesForComponent(component).join('n') ).join('n') + 'n' + 'Write code for making the following:n' + prompt ); // { javascript: '...', css: '...' } |
const components = llm<Array<string>>(
‘You only have the following components: ‘ +
designSystem.getAllExistingComponents().join(‘, ‘) + ‘n’ +
‘What components to do you need to do the following:n’ +
prompt
);
// [‘List’, ‘Card’, ‘ProfilePicture’, ‘TextInput’]
const result = llm<{javascript: string, css: string}>(
‘You only have the following components: ‘ +
components.join(‘,’) + ‘n’ +
‘Here are examples of how to use them:n’ +
components.map(component =>
designSystem.getExamplesForComponent(component).join(‘n’)
).join(‘n’) + ‘n’ +
‘Write code for making the following:n’ +
prompt
);
// { javascript: ‘…’, css: ‘…’ }
What’s pretty magical here is that the llm calls are taking as input arbitrary string but output real values and not just string. In this case it’s using JavaScript and TypeScript for the type definition but can be anything you want (Python, Java, Hack…).
The function that we are using is llm<Type>(prompt: string): Type. It takes an explicit type that will be returned.
The first step is you need to have introspection / code generation from your language to be able to take the type you put and manipulate it. With this type we are going to do two things:
We will convert it to a JSON example and augment the prompt with it. For example in the second invocation the type was {javascript: string, css: string}, we are going to generate: 'You need to respond using JSON that looks like {"javascript": "...", "css": "..."}'. We are using prompt engineering to nudge the LLM to be responding in the format we want.
We also convert it to a JSON Schema that looks something like this:
{ "type": "object", "properties": { "javascript": {"type": "string"}, "css": {"type": "string"} } } |
This is fed to JSONFormer which restricts what the LLM can output to 100% follow the schema. The way LLMs generate the next token is by computing the probability of every single token and then picking the most likely. JSONFormer restricts it to only the tokens that match the schema.
In this case, the first generated tokens can only be {"javascript": " and then the LLM is left filling the blanks until the next " at which point it will be forced to insert ", "css": ", left on its own again and then forced with "}.
The great property of LLMs generating new tokens based on the previous ones is that even without the added prompt engineering, if it sees {"javascript": " it will automatically continue generating JSON and will not be likely to add all the intros like Sure, here is the response.
At this point we are guaranteed to get a valid JSON using our structure. So we can use JSON.parse() on it and then convert it to the JavaScript object we requested.
Before we implemented this magic llm<Type>() function, we’d see people adding a lot brittle logic in order to try and get the LLM to output things in the correct format, do lot of prompt engineering, add fuzzy parsing, retry logic… This was both brittle and added latency to the system.
This is not only a reliability improvement but really unlocks a whole new world of possibilities. You can now leverage LLMs within your codebase to implement functions that returns values just like any other function would, but instead of writing code to run it, you tell it what to do using text.
I’ve been playing Trackmania, a racing game, recently and they introduced a new concept called Cup of the Day. Every day a brand new map is released, for 15 minutes everyone is trying to get the fastest time and based on that time are put into groups of 64 players. Then for 23 rounds people play and the slowest ones are eliminated until there’s only 1 winner left. This is super fun to play!
Each map has 4 times associated: “Author Time”, “Gold Medal”, “Silver Medal”, “Bronze Medal”. Not all those times are of equivalent difficulty for all the maps. I’ve been trying to get gold medals on all the tracks and for some it takes me a few minutes compared to hours for some others. So I’ve been trying to figure out a way to get a sense of how difficult it is to get.
Fortunately, there’s an in-game leaderboard that tells you the times everyone made. So my plan was to scrape this leaderboard, figure out how many people got which medal and hope to get a sense of how hard the map is.

Fortunately, the website trackmania.io has all the information I needed. It has the medal times and the actual leaderboard.

Not only that but the way the website is written is a single page app using Vue that queries the data from a server endpoint using JSON. So this makes retrieving the information even more straightforward, no need to parse HTML.

At this point, what I need is to figure out how to get the number of people that got each medal. The traditional way to do a leaderboard is to have the endpoint return a fixed number of results each time and have a pagination system. So I would do a binary search in order to find where the medal boundary lie.
But, it turns out that this was even easier, the pagination API is doing the limits based on a specific time. So the algorithm was to query the map medal times, and for each of them query the leaderboard and take the position of the first result to know how many people had the previous medal!
For example, Gold time is 1:03.000. I query the leaderboard starting at 1:03, the first person will have 1:03.012 and be at position 2310. They won’t have the gold medal, only silver. But 2309 people will have either the author time or gold medal.
I decided to go with a nodejs script this time around. But you can use whatever language you want, I’ve been doing a lot of scraping in PHP in the past.
The first thing you want to do is to add a caching layer so you don’t spam the server and get you banned, but also make it much quicker to iterate as the next times it’ll be instant. Here’s a quick & dirty way to build caching:
async function fetchCachedJSON(url) { const key = url.replace(/[^a-zA-Z0-9]/g, '-').replace(/[-]+/g, '-'); const cachePath = `cache/${key}.json`; if (fs.existsSync(cachePath)) { return JSON.parse(fs.readFileSync(cachePath)); } const json = await fetchJSON(url); fs.writeFileSync(cachePath, JSON.stringify(json)); return json; } |
And this is what my cache/ folder looks like after it’s been running for a while.

The great aspect about this is that all those are single files that can be looked at manually and edited if needs be. If you someone messed up or got banned, you can delete the specific files and retry later.
If you look at the code, you’ll notice that I didn’t use the fetch API directly but instead used a fetchJSON function. The reason for this is that you’ll most likely want to do some special things.
You probably will need some sort of custom headers for authentication or mime type. It’s also a good place to add a sleep so you don’t spam the server too heavily and get banned.
async function fetchJSON(url) { const response = await fetch(url, { headers: { 'Authorization': 'nadeo_v1 t=' + accessToken, 'Accept': 'application/json', 'Content-Type': 'application/json', 'User-Agent': 'vjeux-totd-medal-ranks', } }); const json = await response.json(); await sleep(5000); return json; } function sleep(ms) { return new Promise(resolve => setTimeout(resolve, ms)); } |
After this, the logic was pretty straightforward, where I would just do the algorithm I described at the beginning. In order to write it, the usual way is to first implement the deepest part and test it standalone (fetching a single time) then wrap it for a map, then for a month, then for all the months.
async function fetchRankFromTime(trackID, time) { const json = await fetchCachedJSON('https://trackmania.io/api/leaderboard/map/' + trackID + '?from=' + time); return json.tops[0].position - 1; } async function fetchRanks(trackID) { const map = await fetchCachedJSON('https://trackmania.io/api/map/' + trackID); const rankAT = await fetchRankFromTime(trackID, map.authorScore); const rankGold = await fetchRankFromTime(trackID, map.goldScore); const rankSilver = await fetchRankFromTime(trackID, map.silverScore); const rankBronze = await fetchRankFromTime(trackID, map.bronzeScore); return [map.authorplayer.name, rankAT, rankGold, rankSilver, rankBronze]; } async function fetchTOTDMonth(month) { const json = await fetchCachedJSON('https://trackmania.io/api/totd/' + month); const days = []; for (jsonDay of json.days) { [authorName, rankAT, rankGold, rankSilver, rankBronze] = await fetchRanks(jsonDay.map.mapUid); days.push({day: jsonDay.monthday, month: json.month, year: json.year, authorName, rankAT, rankGold, rankSilver, rankBronze}); } return days; } async function fetchAll() { const days = []; for (month of [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13]) { days.push(...await fetchTOTDMonth(month)); } console.log('data =', JSON.stringify(days.sort((a, b) => b.rankAT - a.rankAT))); } |
There are two distinct parts of the project. The first part is the data collection, described above. The second is how do you display all this data. I like to keep both separate.
In this case, the end result of the collection is a standalone JSON document where I prepend data =, that I can then include as a <script>. This way in my front-end, also written in Vue for this project, I can use that global data variable to access it.
I didn’t really know how to display the data. We have the number of people that have all the medals before. I started displaying an horizontal bar for each where 1px = 1 position. It worked pretty well but it was way too large.
The Trackmania API stops giving any kind of precision past 10,000 so I used CSS and made all the numbers where 100% is 10,000 and it gave this results which worked well!

Since most tracks are finished by around 8k players it turns out to be working really well in practice.
Now that we have that for every single map, we can start having fun and sort this data in many ways.
Sorting by number of people that got author time gives us what I was looking when going into this project. We can find the easiest maps:

As well as the hardest maps.

I’ve implemented various ways to sort such as date released, medal times and map author name. In doing so, I found out that if 10k people got a medal then the sort is going to give different orderings every time you sort them.
A quick and dirty fix is to sort the data by all the previous pivots so that it always give a stable list. It is wasteful but easy to implement by copy and pasting and the dataset is small enough that it doesn’t matter much in practice.
sortAuthor: function() { this.days.sort((a, b) => (b.year * 10000 + b.month * 100 + b.day) - (a.year * 10000 + a.month * 100 + a.day)); this.days.sort((a, b) => b.rankAT - a.rankAT); this.days.sort((a, b) => b.rankGold - a.rankGold); this.days.sort((a, b) => b.rankSilver - a.rankSilver); this.days.sort((a, b) => b.rankBronze - a.rankBronze); this.days.sort((a, b) => a.authorName.localeCompare(b.authorName)); |
This was a fun project and I’m happy that I was able to figure out how hard a map was in practice. I’d like to give big props to Miss who built all the Trackmania APIs I used during this project.

Early in my time at Facebook I realized the hard way that I couldn’t do everything myself and for things to be sustainable I needed to find ways to work with other people on the problems I cared about.
But how do you do that in practice? There are lots of techniques from various courses, training, tips… In this note I’m going to explain the technique I’m using the most that has been very successful for me: “casting lines”.
So you want something to happen, let say implement a feature in a tool. The first step is to post a message in the feedback group of the tool explaining what the problem is and what you want to happen. It’s fine if it’s your own tool. The objective here is that you have something you can reference when talking to people, you can send them the link with all the context. It can also be an issue on a github project, a quip, a note… the form doesn’t matter as long as you can link to it.
If the thing is already on people’s roadmap or already implemented under a gk, then congratz, you win. But most likely it is not.
This is where you start “casting lines”. The idea is that anytime you chat with someone, whether it is in 1-1 meetings, group conversations, hallway chat… and the topic of discussion comes close (for a very lax definition of close), you want to bring up that specific feature: “It would be so awesome if we could do X”. At you see the reaction. If that person feels interested, you then start to get them excited about them building it. Find ways it connects to their strengths, roadmap, career objectives… and of course send them the link.
In practice, the success rate of this approach, in the moment, is small because people usually don’t have nothing to do right now and can jump on shipping a feature that they never thought about. But if you keep casting lines consistently in all your interactions with people, at some, point someone will bite.
The more lines you cast, the more stuff are going to get done.
While this technique has been very effective at getting things done at scale, there are drawbacks to this approach. The biggest one being uncertainty around timelines. Unless someone bites, you don’t know when something will be done. Some of my lines are still up from many years ago.
PS: while researching for this note, I learned that the fishing technique shown in the cover photo is called “Troll Fishing”.
If you watch pro pool players, most of the time the game is super boring, if you don’t believe me, watch this video from someone that puts 152 balls in a row. What’s interesting is that if you were to look at each shot individually, most of them are easy. I can likely make the 152 pots he did in a row, if I didn’t have to care about positioning myself for the next ball.
The real talent of pro pool players is being able to not only pot the ball but put the white ball in a good position for shooting the next ball. When they play well, they “make the game easy” by having the white ball always in a good position for the next shot.
What this means is that if you see a pro player doing some crazy shot, this means that they “got out of position” in the previous ball. And in practice at this level, usually they did a mistake a few shots earlier and haven’t been able to correct the position back and it gradually amplified.

This is a really bad property when watching the game, so most tournaments introduce a 30s limit so you don’t let the players properly think through and increase the likelihood of making mistakes, and having to come up with interesting shots.
There are a lot of interesting strategies in order to get good at it:
Now is probably the point where you’re asking yourself, that’s interesting but what does it have to do with software engineering. Well, I think that there are a lot of parallels with building software.
When I see people doing very visible and consequential actions, I find myself thinking that they are doing a “hero shot” and it must mean that they got “out of position” for the past few shots and now the only option that they have left is unsatisfying but there’s no other choice.
On the other hand, I see people appearing to somehow always be in easy projects where everything just works out fine and they deliver a lot of impact. I used to think that they were lucky, now I think that they are pro players and are able to plan multiple shots in advance and able to execute on their strategy.
On January 1st I started building a little tool that lets you create diagrams that look like they are hand-written. That whole project exploded and in two weeks it got 12k unique active users, 1.5k stars on github and 26 contributors on github (who produced real code, we don’t have any docs). If you want to play with it, go to Excalidraw.com.
Many people have asked me how I got so many people to contribute in such a short amount of time for Excalidraw, while this is still fresh in my mind, let me post about what I was thinking about during the process.
Before we get started with the actual content, here’s an interesting concept that was in my mind thorough the project. I discovered the concept of a S curve through Kent Beck’s video series. There are three rough phases:
The S curve is usually used to describe bigger projects but it turns out Excalidraw just went through a S curve as seen in this chart that plots the number of stars over the past two weeks.

The most important part for me was to capitalize on the growth phase so that the project doesn’t die when it hits the stabilization phase.
Excalidraw didn’t come out of nowhere, I’ve been using a tool called Zwibbler for probably 10 years in order to build hand-drawn like diagrams to illustrate my blog posts. I’ve always had this feeling that this tool was underrated. I seemingly was the only one to use it even though it felt like it could be used much more broadly.

So when excalidraw came out, there was a clear value proposition and I knew it was going to be somewhat successful. Those days I don’t have that much free time so I tend to spend my time on things that I believe have a high likelihood of being successful, especially side projects.
The first thing was to get people excited! I’m fortunate to have a sizable audience on Twitter so I used it by posting a bunch of videos of the progress of building the first version of the tool.
I got more attention than I anticipated so I felt like I could convert it into actual action. For this, the best way I’ve found is to create a bunch of issues about all the things that need to be done. I’ve been thinking about rebuilding a Zwibbler equivalent for a long time so I had a pretty good sense of what needed to be done.
People that wanted to contribute could just skim through the list of things to be done and start hacking. That worked really well!
When I open sourced React Native, I was convinced that the same people that contributed to React would contribute to React Native. It turns out I was plain wrong, a new set of people started contributing. This same pattern applied to all the subsequent projects I’ve worked on since then.
This is a very broad generalization but most people that tend to contribute significantly to early projects like this are unknown (if they were well known, they’d likely have better opportunities to spend their time) but experienced (they are able to jump in on a random codebase and contribute).
The name of the game is to get as much from people that are interested in contributing as possible. Your initial buzz is only going to last so long (a few days), so you want to capitalize on that time. Everyone (myself included) is likely going to have to go back to their real job soon.
For this, I usually try to be very responsive on the pull requests coming in. If you can get turnaround in less than 10 minutes, then you can have real-time work and people will stay engaged as long as you are.
I’ve tried something new this time and gave commit access to everyone that got a PR merged in. In the past I would do it after I’ve seen sustained work. This worked really well where this gave an extra motivation for people to contribute and they also started to review each other’s code which was awesome! I am not worried about people abusing their power, people that spend energy getting something of quality in tend to be considerate.
A trick I’ve been also using is to merge pull requests even if they’re not exactly the way I want and then push all the follow ups I had in mind. This way the person can have their feature shipped and likely to come back without having expensive back and forth (we never know when / if they’re going to apply suggestions).
People are going to try and stir the project in all sorts of directions with their ideas and pull requests. It’s pretty tricky to think in advance what kind of suggestions you’re going to get because people tend to get very creative (in both good and bad ways…).
If you want something to happen, you need to give a very clear “yes” with concrete things that need to be done. If you’re not sure or change your mind multiple times or answer days/weeks later, people are either not going to invest their time making it happen, or will lose interest and not push it to conclusion.
On the flip side, you’re likely going to see a lot of pull requests or suggestions that you don’t think are a good idea. I’ve found that it’s usually not a good idea to give a clear “no” as it’s a hard message to give to a stranger over text. Instead, what I found tends to work better is to space out replies and ask for more information. The other party will naturally lose interest and move on. You should use this technique very sparingly as it is not a nice approach.
With so many simultaneous contributions, the product can easily start losing quality. I view myself as the keeper of quality. I’ve been pretty obsessed about all the small details and things that feel off.
Every time I see a problem, I open an issue with a small repro case. In many cases, those issues are easy to fix and someone will get to it. I also make sure to clear the backlog so that we’re always in a good enough shape.

I’ve also made sure that some core values were being maintained. I want minimal friction to get started drawing. In particular, this means that what you see first should be the shapes. I had to actively prevent people from adding title selection and login to keep this property.
Posting about all the good things that happen, be it a new cool feature, or interesting usage or thoughts in the topic will increase the size of that channel as those posts will attracts an audience.
The other interesting thing that will happen is that you will provide an audience to a lot of the people that are contributing. As I mentioned earlier, they’re unlikely going to have a big one of their own that cares about this topic.
This is a win-win situation! It takes time to actually post all those things but I’ve seen it being valuable time and time again.
What I found fascinating with this project is that many people were able to project their dreams and ideas onto it. I’ve been told that I should quit my job by at least three people and build a startup around this project as they saw a lot of growth potential in different areas. (Sorry, I’m not, but if you want to, the business is up for grab!)
I’m not exactly sure what to make of that but it led to great conversations! That’s more than I hoped for with this project.
I wish anyone could read this and reproduce it but that’s not completely true. I had a lot of things that went my way. I found it to be useful to know what advantages people behind success stories have to see how they affect their abilities to deliver.
This was a fun project to work on while procrastinating on writing performance reviews. I’m not exactly sure what the future holds for Excalidraw but I’m happy that it is now at a point where I can finally use it to illustrate the blog post I wanted to write that started this whole project (hello rabbit hole!).
Now, go draw some things with excalidraw.com and if you see something you’d like improved, please contribute on github! https://github.com/excalidraw/excalidraw
I’m now [in July 2018] in a group full of compiler engineers at Facebook and learning a lot. Yesterday, I read a post by David Detlefs (summarizing a collaborative idea involving several members of his team) about how to efficiently encode strings for concatenation and since it’s very clever I figured I would share it.
A lot of programs are taking a string as input and building a string as output. You can imagine the following code. Note: I’m going to use JavaScript as an example but it applies to almost all the languages out there.
var str = 'n'; for (var elem of elems) { str += ' * '; if (elem.isExpired) { str += '[expired] '; } str += elem.name + 'n'; } |
An example output might be
'
* Nutella
* Eggs
* [expired] Milk
'
|
What is being executed is
'n' + ' * ' + 'Nutella' + 'n' + ' * ' + 'Eggs' + 'n' + ' * ' + '[expired] ' + 'Milk' + 'n' |
If you implement this naively, the execution would look something like:
'n' + ' * ' = 'n * ' 'n * ' + 'Nutella' = 'n * Nutella' 'n * Nutella' + 'n' = 'n * Nutellan' 'n * Nutellan' + ' * ' = 'n * Nutellan * ' 'n * Nutellan * ' + 'Eggs' = 'n * Nutellan * Eggs' 'n * Nutellan * Eggs' + 'n' = 'n * Nutellan * Eggsn' 'n * Nutellan * Eggsn' + ' * ' = 'n * Nutellan * Eggsn * ' 'n * Nutellan * Eggsn * ' + '[expired] ' = 'n * Nutellan * Eggsn * [expired] ' 'n * Nutellan * Eggsn * [expired] ' + 'Milk' = 'n * Nutellan * Eggsn * [expired] Milk' 'n * Nutellan * Eggsn * [expired] Milk' + 'n' = 'n * Nutellan * Eggsn * [expired] Milkn' |
Because strings are immutable, we need to do a full copy of the string for every small concatenation. In practice this turn a O(n) algorithm into O(n²).
If this becomes a bottleneck, instead of using a string all the way through, you can use an array and push all the string pieces to it. Once you are done building the result, you can join all the pieces together into the final string. Since at this point you know all the strings the operation can sum all the sizes and allocate exactly the right size.
var buffer = ['n']; for (var elem of elems) { buffer.push(' * '); if (elem.isExpired) { buffer.push('[expired] '); } buffer.push(elem.name, 'n'); } str = buffer.join('') |
This pattern works to solve the problem but requires the programmer to know about it and the performance to be bad enough that it is worth writing code in a different way. In practice, a lot of code is not written that way and it’s unclear that any amount of education will change this fact.
Note that a compiler to a bytecode format could, in many cases, make the transformation of the original code to the explicit StringBuffer code. But not in all cases, since compilers have to be conservative: if the string being concatentated is passed as an argument, all bets are off.
The solution that comes to mind is: can normal strings act as a buffer?
The idea is that you allocate a buffer of characters and whenever you do a concatenation, you keep writing at the end of the buffer. If it isn’t big enough, you allocate a bigger one, do a single copy and keep going.
var str = 'n'; // size = 1, ['n', _, _, _, _, _, _, _] str += ' * '; // size = 4, ['n', ' ', '*', ' ', _, _, _, _] // The next one doesn't fit so we need to alloc a new buffer and do a full copy str += 'Nutella'; // size = 11, ['n', ' ', '*', ' ', 'N', 'u', 't', 'e', 'l', 'l', 'a', _, _, _, _, _] str += 'n'; // size = 12, ['n', ' ', '*', ' ', 'N', 'u', 't', 'e', 'l', 'l', 'a', 'n', _, _, _, _] |
// The next one doesn’t fit so we need to alloc a new buffer and do a full copy
str += ‘Nutella’; // size = 11, [‘n’, ‘ ‘, ‘*’, ‘ ‘, ‘N’, ‘u’, ‘t’, ‘e’, ‘l’, ‘l’, ‘a’, _, _, _, _, _]
str += ‘n’; // size = 12, [‘n’, ‘ ‘, ‘*’, ‘ ‘, ‘N’, ‘u’, ‘t’, ‘e’, ‘l’, ‘l’, ‘a’, ‘n’, _, _, _, _]
If you are curious, this is how the Java StringBuilder class is implemented. Performance-wise, this is what we want, but there’s one problem…
You can assign the string to a variable and assign another variable with that variable. For example:
var str = 'n'; var str2 = str; // here we make an alias |
In this case, both str and str2 are pointing to the same 'n' string. In the compiler literature this is called aliasing. The big question is what happens if you try to update one of the variable:
str += ' * '; |
If you look at the JavaScript specification, strings are immutable meaning that you expect str2 to be unchanged but str to be:
str2 == 'n' str == 'n * ' |
Unfortunately, if you mutate the string like in the above solution, then both of them would be 'n * because they both point to the same underlying storage.
If you’ve not been living under a rock, you probably have heard about Rust and linear types. This is a fancy name to say that you cannot have aliasing: there’s only a single variable that can point to a value at all time.
What this means in this case is that the line var str2 = str; would be illegal. If you want to do that, you need to do a full copy of the value so it’s effectively a different one.
In practice, aliasing happens all the time in normal programs, for example calling a function with a string as argument is a form of aliasing. We wouldn’t want to do full copies every time aliasing is happening.
Rust is getting away with it using a concept calling “borrowing” where you can create an alias if the compiler can guarantee that the previous variable cannot be accessed during the lifetime that the alias exists.
In my understanding, you need a strong type system in order to properly enforce those guarantees and in dynamic languages like JavaScript you would have to be too pessimistic and do way more copies than necessary when you are just passing the variable around, ruining the wins you get from building the string in the first place.
Aliasing is usually a dealbreaker because you can mutate the underlying storage that another variable could observe. But in this particular case we can exploit the fact that the only mutation we care about is appending something at the end.
So, in the variable we not only keep a pointer to the buffer but also the size we care about. If someone else appends something at the end, it will not affect us because what’s in the buffer for that size didn’t change.
var str = 'n'; // buffer1, size = 1, ['n', _, _, _, _, _, _, _] // str : size = 1, buffer = buffer1 var str2 = str; // str2: size = 1, buffer = buffer1 str += ' * '; // buffer1, size = 4, ['n', ' ', '*', ' ', _, _, _, _] // str : size = 4, buffer = buffer1 |
var str2 = str;
// str2: size = 1, buffer = buffer1
str += ‘ * ‘;
// buffer1, size = 4, [‘n’, ‘ ‘, ‘*’, ‘ ‘, _, _, _, _]
// str : size = 4, buffer = buffer1
At this point, str2 points to buffer1 with 'n * ' but because it has size = 1 then we know it really is 'n' as intended.
The only edge case to consider is if you are trying to also concatenate str2. If the size of the variable is not equal to the size of the underlying buffer, this means that someone else clobbered the buffer. In this case, our only option is to do a full copy.
str2 += '|'; // buffer2, size = 8, ['n', '|', _, _, _, _, _, _] // str2: size = 2, buffer = buffer2 |
Before joining the team, I knew about the string builder pattern but I had no idea that there was so much theory behind this particular problem like aliasing, linear types… I hope that explaining those concepts in terms of JavaScript is helpful to get some insights into what’s happening inside of compilers.
Andres Suarez pointed me to some interesting code in the Hack codebase:
let slash_escaped_string_of_path path = let buf = Buffer.create (String.length path) in String.iter (fun ch -> match ch with | '\' -> Buffer.add_string buf "zB" | ':' -> Buffer.add_string buf "zC" | '/' -> Buffer.add_string buf "zS" | 'x00' -> Buffer.add_string buf "z0" | 'z' -> Buffer.add_string buf "zZ" | _ -> Buffer.add_char buf ch ) path; Buffer.contents buf |
What it does is to turn all the occurrences of , :, /, and z into zB, zC, zS, z0 and zZ. This way, there won’t be any of those characters in the original string which are probably invalid in the context where that string is transported. But you still have a way to get them back by transforming all the z-sequences back to their original form.
The first interesting aspect about it is that it’s using z as an escape character instead of the usual . In practice, it’s less likely for a string to contain a z rather than a so we have to escape less often.
But the big wins are coming when escaping multiple times. In the escape sequence, it looks something like this:
-> \ -> \\ -> \\\\ -> \\\\\\\\whereas with the z escape sequence:
z -> zZ -> zZZ -> zZZZ -> zZZZZThe fact that escaping a second time doubles the number of escape characters is problematic in practice. I was working on a project once where we found out that the character represented 70% of the payload!
It’s way too late to change all the existing programming languages to use a different way to escape characters but if you have the opportunity to design an escape sequence, know that escape sequence is not always the best 🙂
It’s very trendy to bash at Computer Science degrees saying that it costs a lot of time and money and at the end, you haven’t learned much useful things for your day job. My experience going to EPITA, a French school in Paris, has been the complete opposite!
I started programming when I was around 10. By 13, I was already being contracted for real money by someone in the US (Hi Thott!). So, when I started EPITA at 18, I already had a ton of experience as a self-taught programmer. My biggest fear was: “Would I learn something new?” and, I did learn a ton! I know wouldn’t have done all the impactful things I have at Facebook without it.
What struck me was that I already knew a lot and I knew there was a lot left to learn, but I didn’t really know what nor how to get started learning it on my own.
Here are some themes where I learned entirely new domains of knowledge at school, many of which I would likely not have learned, or at least not as in depth if I didn’t do a CS degree.
What was awesome about EPITA is that not only we had classes like any normal school but also a ton of actual non-trivial projects to implement. Here are some that had a profound impact on me:
And many more… but that’s a good enough list for now!
Now, the big question is: but did you use any of that at work? And the answer is yes, so many times!
next: page in the markdown and not have to have to write prev nor have a global index and yet have an table of content. It took me a second to realize that I had a list of edges of a graph and I needed to rebuild a graph from it. Because I’ve written so many graph traversals at school, it took me under an hour to build it correctly.I could have gotten a well paying programming job right out of high school as a web developer, but going to EPITA made me such a better programmer because I’ve been exposed in depth to a lot of areas that I would likely have never gone on my own.
Now, if I see a problem on a web developer task, I also know that changing the programming language, using machine learning, writing code in a more C-like way, using other data structures and algorithms… are possible tools to solve that problem. More importantly, if I need to use those, I can implement them because I’ve already done similar things in the past.
Your experience may vary, but for me, getting a CS degree at EPITA was a really good decision and I would do it again if I had to.