Over the past few months, I have submitted quite a few PRs to hikyuu, a high-performance C++ quantitative trading framework on GitHub. Along the way, I picked up some useful, reusable experience, so I wanted to put it all together here.

All About the Repo#

hikyuu is a huge, high-performance quantitative backtesting framework. It contains not only C++, but Python as well. If you rely on a human alone to optimize it and contribute to it, understanding the whole framework is extremely difficult.

I am fairly interested in quantitative trading, which is how I stumbled across the framework. For me, stability, efficiency, and continuous maintenance are the three things that determine whether I will actually use a framework.

However, because quantitative backtesting has some rather special requirements, I need to know that the framework I use is under my control and, more importantly, correct. So I used the now-very-common Coding Agents to run a deep, detailed scan of the entire repo. They uncovered quite a few correctness-related bugs, and that was how I started fixing them.

The Development Environment#

Setting Things Up#

I had not worked much with C++ environments before. I still had Visual Studio 2022 installed from a C++ class at school, so naturally, I reused it. Since I did not have GCC installed locally, I used the MSVC toolchain from VS 2022 as the compilation environment for this project.

I created a dedicated Git branch for adapting the build to my local machine.

Choosing the Models#

Throughout the development process, one model I particularly enjoyed using was Zhipu’s GLM 5.2, hosted on Ollama Cloud. At the time, I also used Kimi K3 through Kimi’s official Coding Plan for the broader development and repair work. GPT 5.5 in Codex also handled parts of the fixes and PR submissions.

Meanwhile, I used Gemini 3.1 Pro Preview for structural audits and quality control.

Choosing the Harness#

For this project, I mainly used Zhipu’s ZCode as the Harness environment, with Codex taking on part of the work as well.


C++ has never received all that much attention in either model development or benchmark design. DeepSWE, for example, does not include a single C++ case in its benchmark. That makes workflow and structural optimization especially important for Agentic Coding.

None of the 113 tasks in DeepSWE v1.1 uses C++

DeepSWE v1.1113 tasks · snapshot from 2026-09-11
1 square = 1 taskselected language
C++
0 / 113
Share of sample: 0% No matching tasks
Choose a language to see where it appears in the sample.
Data source and counting method
  • Data snapshot: 2026-09-11, DeepSWE v1.1, from locally recorded commit 0b9fabb.
  • Language distribution: 35 Go, 34 TypeScript, 34 Python, 5 Rust, and 5 JavaScript, for 113 tasks in total; C++ has 0.

Optimizing the Agentic Coding Workflow#

Modern models are generally MoE models, with a huge number of total parameters but a relatively small number of active parameters. In other words, an MoE model activates a different set of parameters on each run, so its performance is often less stable than we might expect. At the same time, we have also seen that nearly every model can perform extremely well on one particular benchmark.

That leads to a question: how can we activate more of the model’s ability—and activate the right parts of it?

Divide and Conquer: Activate the Right Parameters#

In 2025 and earlier, most model training placed a great deal of emphasis on mathematical and logical ability. Starting in 2026, models gradually began receiving more training for Agent capabilities. In other words, the models on the market now have both extremely strong reasoning ability and extremely strong long-horizon Agent capabilities. So why, when we ask them to write code, do they so often fail to execute correctly or show the excellent reasoning ability we know they have?

The answer is that those two capabilities are not perfectly fused. Humans are much the same: trying to think and talk at the same time is actually very difficult.

So how do we genuinely make fuller use of both capabilities?

Simple: split the task.

Separate the work aimed at logic and mathematical reasoning from the work aimed at actually executing the code changes.

Take a backend development task as an example. We are facing a tangled, absolutely enormous codebase, and somewhere inside it is an extremely stubborn, very old bug that we need to change. What would we normally do? Right: trace the call-chain logic, make a plan, modify the code, and run the tests.

Writing code with AI should work the same way.

Here is what I do:

  1. Map the entire codebase in depth. Find the state-machine call chain relevant to the task, turn it into a mathematical abstraction, and display it as an ASCII diagram.
  2. Perform a deep mathematical and topological analysis of that abstract state machine and call chain. Identify conflicts, contradictions, and bugs, then produce a report.
  3. Have a SubAgent deeply review all code related to the task, untangle every potentially relevant contextual call chain, and produce a practical, executable plan.
  4. Execute the plan and complete all tests.
  5. Use a SubAgent to meaningfully review the tests and audit the quality of the code.

A five-step workflow that separates reasoning, execution, and review, with explicit artifacts handed from one step to the next

Reasoning · 1—3Execution · 4Review · 5
  1. Relevant code → state-machine and call-chain model
  2. Model analysis → conflict points and bug report
  3. SubAgent rereads the context → executable plan
  4. Implement the plan → code patch and test results
  5. SubAgent reviews tests and code → audit conclusions
InputTask-related code and call chain
ArtifactState-machine + call-chain model
1 / 5 · Reasoning

That is it. It is an incredibly, incredibly simple setup, but it can let DeepSeek V4.1 Flash beat GPT-6 Astra outright. (Yes, this is exactly the trick I used before when helping a friend untangle a codebase and find the right bug fix. It works every time.)

That said, this experience is mostly limited to certain backend tasks. As for frontend work, well, that is something everyone will have to explore for themselves….

Conclusion#

As Agentic Coding continues to develop and models become more capable, I think more and more people will realize just how much good engineering can amplify what a model is able to do.

Hope you all have fun.

Happy coding!