@brickpage65
Profile
Registered: 4 years, 1 month ago
You Should Try To Process Them! Sloc Cloc and Code, a command line tool I developed by Sloc Cloc and Code, (which is now being modified and maintained in many other excellent ways) counts lines and comments and estimates the complexity of files in a directory. You need to have a good sample size in order to make good use the latter. The way it works is that it counts branch statements in code. But what does it actually mean? Without context, the statement "This file has 10 complexity" isn't very useful. I thought it would be a good idea to run scc through all the source codes I could find. This would allow me see if there were any edge cases I didn’t consider in the tool. A brute force Q/A trial by fire. However if I am going to run it over all that code, which is going to be computationally expensive I may as well try to get some use out of it. I decided to take notes as I went along and see what I could come up with. In short I downloaded and processed a lot of code using scc. The raw numbers include: 9,985,051 total repositories 9,100,000.083 repositories having at least one identified file 884.968 empty repositories, (those without files) 3,529.516,251 files in all repositorys 40.736,530.379,778 bytes processing (40 TB). 1,086,723,618,560 code lines identified 816.822,273,469 code line identified 124.382,152.510 blank lines identified. 145,519,192.581 comment lines identified. Lets get the elephant out of the room first. It wasn't 10 million projects, as the "click bait” title suggests. I was shy by 15,000 so I rounded up. Please forgive me. It took about 5 weeks to download and run scc over the collection of repositories saving all of the data. It took just under 49 hours to crunch the 1TB of JSON data and produce the below results. Also note its likely I messed up some of the calculations. As soon as possible, I will keep this updated with the data set. Quicklinks Methodology Presenting and Computing Results Cost Data Sources How many files in a repository? How is the project split by language? What is the average number of files per language in a repository? How many lines of code are in a typical file per language? Average complexity for file in each language? Average comments for file in each languages? What are the most used filenames How many repositories are missing a license? Which languages have the most comments? How many projects use multiple .gitignore files? Which language developers have the biggest potty mouth? Longest files by lines per language Whats the most complex file in each language? What is the longest file, ranked by number of lines? Whats the most commented file in each language? How many "pure” projects are there? Projects that use TypeScript, but not JavaScript. CoffeeScript and TypeScript users? How long is the typical path, broken up by language (YAML/YML)? Mixed case, lower or upper case? Java Factories Do not bother with files Future ideas Raw / processed Files Methodology Searchcode.com already has over 7,000,000 projects in git, subversion, and mercurial. I thought it would be fun to try processing them. Working with Git is usually the easiest way to do this so I ignored subversion mercurial and instead exported the full list. It turned out that I actually had 12,000,000 git repositories being monitored, so I should probably update this page. So now I have 12,000,000 or so git repositories I need to download. You can have scc output the results in JSON by running scc --format JSON --output myfile.json.go. The results will look like this (for a single file). Here are the JSON results for redis.json. This output contains all the results, without any additional data. Keep in mind that scc categorizes languages based on their extension (except for extensions that are shared like Verilog or Coq). As such, a HTML file containing a Java extension will be counted to be a J2EE file. This is not usually a problem because you wouldn't do that. It is, however at large scales. It is something I discovered later, where files were masquerading in another. A while back I wrote code to create github badges using scc https://boyter.org/posts/sloc-cloc-code-badges/ and since part of that included caching the results, I modified it slightly to cache the results as JSON in AWS S3. I used the badge code to work in AWS using Lambda. I exported the list of projects and wrote about 15 lines python. I threw in some python multiprocessing to fork 32 processes to call the endpoint reasonably quickly. This worked wonderfully. However the problem with the above was firstly the cost, and secondly lambda behind API-Gateway/ALB has a 30 second timeout, so it couldn't process large repositories fast enough. Although I knew this wasn't the most cost-effective solution, I was willing to accept it if it was close to $100. After processing 1 million repositories I checked and the cost was about $60 and since I didn't want a $700 AWS bill I decided to rethink my solution. Keep in mind that this cost was primarily storage and CPU. This was what was needed to collect the information. Assuming that I processed the data or exported it, it was going up in cost significantly. The best solution for me would be to dump the URL's as messages into SQS, and then pull from it using either EC2 instances or fargate. This will allow you to scale up like mad. However despite working in AWS in my day job I have always believed in taco bell programming. It was only 12,000,000 repositories, so I chose to implement a simpler (and cheaper) solution. Running this computation locally was out due to the abysmal state of the internet in Australia. However I do run searchcode.com fairly lean using dedicated servers from Hetzner. These boxes are quite powerful. They often have 2 TB of storage space, which is usually unused, and are i7 Quad core 32 GB RAM machines. As such they usually have a lot of spare compute based on how I use them. The front-end varnishbox, for example, does the square root and zero most of all the time. Why not do the processing there? I didn't quite taco bell program the solution using bash and gnu tools. I wrote a Go program that spun up 32 go-routines. These read from a channel and then spawned git, scc subprocesses. Finally, the JSON output was written into S3. I did actually create a Python solution initially, but it seemed like a bad idea to have to install the pip dependencies upon my clean varnishbox. It kept breaking and I didn't feel like debugging. Running this on the box produced the following sort of metrics in htop, and the multiple git/scc processes running (scc is not visible in this screen capture) suggested that everything was working as expected, which I confirmed by looking at the results in S3. Game servers Computing and presentation of results Having recently read https://mattwarren.org/2017/10/12/Analysing-C-code-on-GitHub-with-BigQuery/ and https://psuter.net/2019/07/07/z-index I thought I would steal the format of those posts with regards to how I wanted to present the information. My twist on the previous however is to add jQuery DataTables over the large tables of information. This allows you sort and filter results. You can click the headers for sorting and the search field to filter. The search box indicates that this is enabled for you. I have also added a jump button near these tables so that you can skip over them if necessary. The size of the data I needed to process raised another question. How can you process 10 million JSON files that take up just over 1TB of disk space in an S3 bucket. The first thought I had was AWS Athena. It's going to be $2.50 USD for each query for that dataset, so I quickly searched for an alternative. That said if you kept the data there and processed it infrequently this might still work out to be the cheapest solution. I posted the question on company slack because I don't believe I should solve all problems alone. One suggestion was to dump the data into large SQL databases. This would require processing the data into the SQL database and then running multiple queries. Because of the data structure, there were a few tables that required foreign keys and indexes in order to provide some performance. This feels wasteful because we could just process the data as we read it from disk in a single pass. It was also a concern to build a database this big. It would take more than 1 TB of data to build it, before adding indexes. I thought it would be a good idea to process the JSON the same way as I did the spare compute. This has one problem. Pulling 1 TB of data out of S3 is going to cost a lot. In the event the program crashes that is going to be annoying. I wanted to save the files locally to cut down on costs. Handy tip, you really do not want to store lots of little files on disk in a single directory. It is a waste of time and a disaster for file-systems. I used another program called go to pull down the files from S3 and store them in an tar file. I could then process this file over and over again. The process itself is not very pleasant. I used a tar program to extract the tar file. This allowed me to run my questions again, without having S3 to search for them. Two reasons prompted me to not use go-routines with this code: First, I didn't want my server to be overloaded so this limits it only to one core for hard CPU work (another core to read the tar files was mostly blocked on the processor). The second reason was that I didn’t want it to be thread-safe. Now I needed a bunch of questions to answer. I again used the slack Brains Trust and crowdsourced my colleagues at work while I created some of my own ideas. Below is the end result of this mind meld. You can find all the code I used to process the JSON including that which pulled it down locally and the ugly python script I used to mangle it into something useful for this post https://github.com/boyter/scc-data Please don't comment on it, I know the code is ugly and it is something I wrote as a throwaway as I am unlikely to ever look at it again. You can review the code I have written for others by going to the source of scc. Cost While I was testing lambda, I spent $60 on compute. I haven't looked at the S3 storage costs yet, but it should be close $25 based upon the size of my data. However, this does NOT include the transfer fees which I have also not noticed. It is important to note that I cleared the bucket after I was done with it, so this is not an ongoing expense for me. AWS was expensive so I decided to stop using it in the end. What's the actual cost, assuming I want to do it again? Well to start all the software used is free as in freedom and open source. So nothing to worry about there. In my case, the cost was free because I used "free compute" leftover from searchcode. However, not everyone has compute. Let's suppose that another person wants to do the same thing and needs a server. It could be done for EUR73 using the cheapest new dedicated server from Hetzner https://www.hetzner.com/dedicated-rootserver However that cost includes a new server setup fee. If you are willing to wait and poke around on their auction house https://www.hetzner.com/sb you can find much cheaper servers with no setup fee at all. I found the following machine, which is EUR25.21 per monthly and has no setup fee. The best part is? You can get the VAT eliminated if you're not in the EU. So give yourself an additional 10% discount on top if you are in this situation. If you were to do it from scratch using the exact same method I used, it would likely cost under $100 USD to redo these calculations. If you are patient or lucky, it might be less than $50. This assumes that you only use the server for a maximum of 2 months, which is enough time to download and process. This allows you to download a list from 10 million repositories that you can then consider processing. If I were to use the gzipped.tar.tar file for my analysis (which isn’t very difficult to do really), I could do 10x the repositories from the same machine. However, the resulting file would still be small enough that it could fit on the same disk. However, this would take longer and increase the cost per month. Additionally, it could take multiple months to download. Going much larger then 100 million repositories however is going to require some level of sharding. You can still do the same process as I did, or a larger one, on the same hardware. It doesn't take much effort to make changes or make code changes. Data Sources How many projects came from each of these three sources, bitbucket gitlab and github. This is done before counting empty repositories, so the sum is greater than the number of repositories which actually make up the counts below this point. Sorry to the GitHub/Bitbucket/GitLab teams if you read this. If you find this offensive (which I doubt), I will give you a refreshing beverage of choice if we ever get together. How many files in a repository? Now let's get to work. Let's start with a very simple question. How many files are in an average repository? Are there a lot of files or fewer? By looping over the repositories and counting the number of files we can then drop them in buckets of 1, 2, 10, 12 or however many files it has and plot it out. In this example, the X-axis is a bucket of the total number of files and the Y-axis is the count of projects that have those many files. This is limited to projects with less than 1000 files because the plot looks like empty with a thin smear on the left side if you include all the outliers. It turns out that most repositories contain less than 200 files. However what about plotting this by percentile, or more specifically by 95th percentile so its actually worth looking at? The vast majority of projects have less that 1,000 files. 90% of them have less then 300 files, while 85% have less the 200 mark. If you want to plot this yourself and do a better job than I here is a link to the raw data filesPerProject.json. Whats the project breakdown per language? This means that each project scanned will have its Java count incremented by one for each Java file identified. For the second file, it will do nothing. This gives a quick view of what languages are most commonly used. Not surprisingly, markdown, plain text and.gitignore are the most popular languages. Only about 2/3rd of all projects have Markdown as the most popular language. This is understandable as almost all projects have a README.md that is displayed in HTML for repository pages. Below is the complete list. How many files are in a repository per language An extension of the previous, but averaged over however much files are in each repository's language. So, for projects that include Java, how large are the java files in that project and, on average, how big are all the files? This will allow you to see if a particular project is larger or smaller than expected for your language. How many lines of code can be found in a typical language file? This could also be used to determine which languages have the largest file sizes. Using the average/mean for this pushes the results out to stupidly high numbers. This is because projects such as sqlite.c which is included in many projects is joined from many files into one, but nobody ever works on that single large file (I hope! ). This was then calculated using the median. There are still definitions that have ridiculously high numbers, such as JavaScript and Bosque. So I thought why not have both? I did one small change based on the suggestion of Darrell (Kablamo's resident and most excellent data scientist) and modified the average value to ignore files over 5000 lines to remove the outliers. Average file complexity in each language What is the average complexity of each file? The complexity estimate isn't really directly comparable between languages. The README in scc is helpful. Complexity estimates are only comparable to files written in the same language. It should not be used to directly compare languages without weighting them. This is because it calculates by looking for loop and branch statements in the code, and incrementing a counter to that file. It is not a good idea to compare languages, even though it might be possible to have some similarities between languages like Java and C. Your mileage may vary. This is especially useful when it applies only to files that are part of the same language. You could also ask the question "Is the average file I work with more or less complicated than the average?" I should also mention that this calculation is constantly improving and I am looking for submissions from scc. Usually, it is as simple adding keywords to the languages.json so any programmer of whatever skill level should be capable of helping. Average comments for file in each language? What is the average number of comments per file in each language? If you're able to see enough, you could probably rephrase the question to ask developers which developers create the most comments. What are the most used filenames? What filenames are most commonly used across all code-bases, ignoring extension or case? Had you asked me before I started this I would have said, README, main, index, license. Thankfully the results reflect my thoughts pretty well. Although there are a lot of interesting ones in there. I don't know why so many projects have a file called 15/s15. I was surprised that the makefile was the most used. But then, I remembered that it is used by many JavaScript projects. Another interesting fact is that jQuery seems to still be the king. Reports of its death are greatly exaggerated. It appears as #4 on our list. This was due to memory constraints. Every 100 projects I checked I would check my map. If a filename was identified as having 10 at this point, it would remain. This shouldn't happen often, but it is possible to have some errors in the counts if some common names appeared sparsely within the first batch. In short they are not absolute numbers but should be close enough. I could have used a trie structure to "compress" the space and gotten absolute numbers for this, but I didn't feel like writing one and just abused the map slightly to save enough memory and achieve my goal. However, I am curious enough to test this at a later time to see how a triangle would perform. How many repositories are missing a license? This is a very interesting question. Which repositories do not have an explicit license folder? It is important to note that a project might not have a license file if it does not exist in the README. It is simply that sccc couldn't locate an explicit license folder using its own criteria. This, at the time of writing, means that it ignores cases such as "license", 'licence", 'copying", copying3", unlicense", licence",?licence-mit",?licence–mit" or?copyright. Sadly it appears that the vast majority of repositories are missing a license. Although I strongly believe that software should have a license, here is another viewpoint. How many projects use multiple.gitignore file? This is a fact that many people don't know, but it is possible for a git repository to contain multiple.gitignore directories. Given this fact, how many projects have multiple.gitignore folders? While we are looking how many have none? I was intrigued to find that one project has 25,794.gitignore folders in its repository. The next highest was 2,547. I don't know what's going on. I took a quick look at it, and it appears they are used to allow checking into directories. However, I am unable to confirm this. Here is a plot of data up to 20.gitignore file and close to 99% for the total result. Something you would expect would be that the majority of projects would have either 0 or 1 .gitignore files. The results confirm this with a dramatic drop-off in 10x for projects that have 2.gitignores. Surprisingly, many projects have more than one.gitignore. This is a case where the tail is especially long. I was also curious to know why some projects had thousands.gitignore folders. One of the main offenders appears to be forks of https://github.com/PhantomX/slackbuilds which all have ~2,547 .gitignore files. Below are the links to other repositories that have more than 1000 ignore files. - 25,794 files https://github.com/knot-git/rrs_publication_ocr - 1,523 files https://github.com/brunosimon/hetic - 1,267 files https://github.com/par4all/par4all - 1,186 files https://github.com/ramonluis1987/code_school - 1,133 files https://github.com/signed/license-maven-plugin Which language developers have the biggest potty mouth? Working this out is not an exact science. This falls under the NLP category of problems. Picking up offensive terms or cursing/swearing from files in a list is not going to work. If you run a string contains test, you will find all kinds of files like assemble.sh. So to produce the following I pulled a list of curse words, then checked if any files in each project start with one of those values followed by a period. This would mean a file named gangbang.java would be picked up while assemble.sh would not. This is however going to miss all sorts cases like pu55syg4rgle.java, and other such crude names. To try to catch the most interesting cases, I used some leet talk such as b00bs or b1tch in the list. You can find the complete list here. Although this is not accurate, it is still quite entertaining to see what the result of this experiment produces. Let's start by listing the languages with the most curse words. This should be weighed against the amount of code. These are the top ones. Interesting! It was my first thought to think, "those naughty C-developers!" But it turns out that while they have a large number of developers, it isn't so big a deal. Dart developers clearly have an axe. You may want hug someone who codes with Dart. I also want to know what are the most commonly used curse words. Let's see what kind of minds we all have. A few of the top ones I could see being legitimate names (if you squint), but the majority would certainly produce few comments in a PR and a raised eyebrow. Note that some of the more offensive words in the list did have matching filenames which I find rather shocking considering what they were. Thankfully they were not very common and didn't make my list above which was limited to those which had counts over 100. I hope that these files are only used for testing allow/deny lists, and such. Longest files by lines per language Plain Text, SQL and XML are the top positions. CSV is in the middle. Which file is the most complicated in each language's language? These values are not directly comparable, but it's interesting to see which language has the most complex. Some of these files may be absolute monsters. For example consider the most complex C++ file I found COLLADASaxFWLColladaParserAutoGen15PrivateValidation.cpp which is 28.3 MB of compiler hell (and thankfully appears to be generated). What is the most complex file weighed against lines? This sounds great in theory but it is a waste of time. As such, I did not include this calculation. I have however created an issue inside scc to support detection of minified code so it can be removed from the calculation results https://github.com/boyter/scc/issues/91 It is possible to infer the above using only the data at hand. However, I would like to make it a more robust test that anyone using scc could benefit from. Whats the most commented file in each language? Whats the most commented file in each language? Although I don't know what kind of information you might get from this, it is worth a look. NB: Some links below may not translate 100% because I lost some information when creating the files. Most should work, but a few you may need to mangle the URL to resolve. How many "pure" projects Assuming pure is one project with 1 language. It would be quite boring, so let's find out what the spread is. It turns out that most projects have less than 25 languages, with the majority in the 10 to 10 range. The graph below shows 4 languages. Pure projects might only use one programming language, but may also have many supporting formats such as css, markdown, and yml which are picked by scc. It's reasonable to assume that projects with less than 5 languages are "pure" for some level of purity. This is based on the fact that they have just over half of the total data sets. You can adjust the number to fit your own definition of purity. What suprises me is an odd bump around 34-35 languages. I don't know why this is the case, and it probably warrants some investigation. Below is the complete list. TypeScript and not JavaScript are used in projects The modern world of TypeScript! For projects that use TypeScipt, how many are exclusively using TypeScript? This number surprised me a bit. Although mixing JavaScript and TypeScript is quite common, I would expect there to be more projects using this new hotness. This could be due to the projects I was able pull through. I suspect that a refreshed project list with more recent projects would change this number dramatically. CoffeeScript or TypeScript are you using? I get the feeling that some TypeScript developers are shivering at the mere thought of this. If it is of any comfort I suspect most of these projects are things like scc which uses examples of all languages mixed together for testing purposes. What is the average length of a path, broken down by language? Given that you can either dump all the files you need in a single directory, or span them out using file paths whats the typical path length and number of directories? This is done by counting the number of path separators / for each file and its location and averaging it out. I didn’t expect much other than Java being near the top, as its file path are often quite deep. YAML or YML Sometime back on the company slack there was a "discussion" with many dying on one hill or the other over the use of .yaml or .yml The debate can now (?) The debate can (????) be ended. Although I suspect others will still prefer the hill to their death. Upper lower or mixed case? What case style is used for filenames? This includes the extension, so you would expect it mostly to be mixed case. This is not interesting, as file extensions are usually lowercase. What if we forget the file extension Not what I would have expected. Mixture is normal, but I would expect lower to be more popular. Java Factories Another one was discovered in an internal company slack while looking through old Java code. I thought why not add a check for any Java code that has Factory, FactoryFactory or FactoryFactoryFactory in the name. The idea being to see how many factories are out there. So slightly over 2% of all the Java code that I checked appeared to be a factory or factoryfactory. Thankfully there are no factoryfactoryfactories and perhaps that joke can finally die, although I am sure at least one non-ironic one exist somewhere in some Java 5 monolith that makes more money every day than I will see over my entire working life. Do not bother with files The .ignore file idea was hammered out by burntsushi and ggreer in a Hacker News thread and is possibly one of the greatest cases of "competing" open source tools working together to a good outcome and done in record time. It is the de facto way to add things into source controls, yet tools ignore them. As it turns out, sccc not only implements.ignore but counts them too. Let's see if the idea spreads. Skip to next section Future ideas Id love to do some analysis of tabs vs spaces. A scan for AWS AKIA Keys and the like would also be very useful. I would love to see more bitbucket/gitlab coverage, and have it broken down through each to see where developers from different camps congregate. If I'm able to do it again, there will be some shortcomings that I would love to fix. - Properly storing the URL in metadata. This was a bad idea. A filename is a lossy way to store it. It can be difficult to find the file source and location. - Not bother with S3. If I was only using the bandwidth for storage, it is no point paying bandwidth fees. Better to just stuff into the tar file from the beginning. - Spend some time learning a tool to help you plot and chart your results. - Instead of using the less accurate approach I used, you can use a trie (or another data type) to keep a complete count. - Add an option to scc to check the type of the file based on keywords as examples such as https://bitbucket.org/abellnets/hrossparser/src/master/xml_files/CIDE.C was picked up as being a C file despite obviously being HTML when the content is inspected. To be fair, all of the code counters I tried behaved in the same manner. - It appears there is a bug with sccc that if a file does not have an extension, but is named as one, it will match that file. A bug has been raised in scc to address this https://github.com/boyter/scc/issues/114 - I'd like to add shebang detection into scc https://github.com/boyter/scc/issues/115 - Some sort of check against number of github stars would be pretty neat. - An analysis of the number commits would be very useful. - I want to add maintainability index calculations at some point. It would be very cool to see what projects are considered the most maintainable based on their size. So why bother? You can use some of this information to plug it into searchcode.com or scc. Even if only some useful data points. This was the stated goal and it can be very useful to see how your project compares with others. It was also a fun and interesting way to spend some time solving interesting problems. It is also quite reliable, I believe. I am also working on a tool that will help senior-developers and managers analyze code looking for flaws, large file sizes, languages, etc... with an assumption that you must monitor multiple repositories. It will take some code and tell you how manageable it is. It can help you determine if you need to buy or maintain a code-base, and give you a snapshot of the work done by your development team. This could help teams scale with shared resources. This is how I see it. It's something that I use for my day job, and I think others might find it useful. To get more interest in this, I would probably put an email signup here. Raw / processed files I have provided a link (20 MB) to the processed files for anyone who would like to do their own analysis or corrections. If you would like to share the raw files with others, let me know. It is a 83GB tar.gz file that uncompressed is just above 1 TB in size. It contains just over 9,000,000 JSON files in various sizes.
Website: https://riber-anderson.blogbright.net/on-top-of-this
Forums
Topics Started: 0
Replies Created: 0
Forum Role: Participant