Hello! Here is for your evaluation a patch I made to improve the linker performance. In my project with 61705 symbols it was taking ~26 seconds to link but now it links in under 2 seconds. The gist of it is that I changed the current separate chaining hashmap with boost::unordered_flat_map and added a map of sets for each area to keep track of which symbol is in which area that avoids the terrible double loops that are used in lstarea. I don't think the patch is ready to be merged as it is (I haven't...
I'll take a look! How do I run the whole regression suite to validate my changes?
I made a bit of progress running some tests with Valgrind. For gcc-torture-execute-strlen-4.c I went from time ../../bin/sdcc --fverbose-asm -DNO_VARARGS --nostdinc -I../.. -msm83 -c 216.43s user 0.34s system 99% cpu 3:36.87 total to time ../../bin/sdcc --fverbose-asm -DNO_VARARGS --nostdinc -I../.. -msm83 -c 127.92s user 0.40s system 99% cpu 2:08.42 total By observing that this bit of code in cseBBlock: for (expr=setFirstItem (ebb->inExprs); expr; expr=setNextItem (ebb->inExprs)) if (!isinSetWith...
Sorry if I'm wasting your time with this questions, but I'm genuinely trying to learn and hopefully help. Is it know why this implementation is problematic in contrast with other compilers that doesn't seem to have the same problem? Would SSA form help?
Correct me if I'm wrong, but from my limited testing it seems that cseBBlock is currently a big bottleneck for general compilation. I'm genuinely interested in dedicating a bit of time to try to optimize it, even if it's not general but adding flag that can make it recurse a bit less for non-optimized builds. I'm not an expert but I can learn. Does it make sense this approach, or is there any other structural issue that would prevent focusing on this section of code from getting improvements? Any...
This is becoming an actual bottleneck in my project. Having to wait minutes for a change in one file is killing productivity and tracking/finding which functions might be responsible is not easy and time consuming. In cseAllBlocks a loop calls cseBBlock for every block. Do you think this calls could be done in different threads at the same time? (I'm willing to try at least for a local patch to speed things up).
Hello! Now that my project is getting bigger I'm finding that some files sometimes cross a threshold where compilation changes from a couple of seconds to minutes. Even using quite low max-allocs-per-nodesettings. I have no idea why it happens. Is there any way to profile the compiling process to at least pinpoint where it's spending it's time? If it's just a tradeoff of how the compiling/optimization works I have no problem splitting functions or making changes to code to make it easier for the...
An array of pointers indeed works but I wanted to avoid the extra indirection when accessing the far pointers. I can check the linker support trying to build the array in asm. That would work for arrays but I'll still have the same problems for structs initialization.