Key Topics:
SWE Benchmarks for LLMs and Agents:
Assessing code quality (linting, style compliance).
Success rates on unit tests or example test suites.
Accuracy metrics or correctness checks.
Tools like SWE Lancer or internal test harnesses.
Interpreting Results:
Understanding false positives or incomplete suggestions.
Distinguishing between speed vs. correctness vs. maintainability.
Making Adoption Decisions:
Weighing cost, team familiarity, ecosystem integration, and performance.
Potential fallback or layering strategy (e.g., default to one assistant, then cross-check with another)
Below you will find some suggested materials on these topics, which we believe may help you complete the final activity in the lesson. Please note that using these resources is entirely optional, and you are welcome to explore any other sources or courses at your own discretion.
Here are some recommended courses to explore:
Key Topics:
SWE Benchmarks for LLMs and Agents:
Assessing code quality (linting, style compliance).
Success rates on unit tests or example test suites.
Accuracy metrics or correctness checks.
Tools like SWE Lancer or internal test harnesses.
SWE Lancer GitHub Repository
An open-source project focusing on standardized test suites for AI-generated code.
Explore issues/discussions to understand how devs measure pass/fail rates and style compliance.
Community-driven posts showing real-world approaches to testing code generated by LLMs. Often includes scripts or examples for measuring correctness vs. code style.
SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at $1 million USD total in real-world payouts.
2025 -Learn how Gen AI , AI Agents, AI plugins can boost your QA Automation productivity & demo of AI Powered Test tools
For classifier models, “Accuracy” is defined: Correctness on the intended task is typically the primary consideration for a model developer
Key Topics:
2. Interpreting Results:
Understanding false positives or incomplete suggestions.
Distinguishing between speed vs. correctness vs. maintainability.
Artificial intelligence (AI) can be used not only in software development, but also in software testing. One of the use cases is also using AI in the analysis of automation tests results.
Search “AI false positives code suggestions” to find dev experiences with AI guesswork.
Many discuss how they weigh speed vs. final correctness or maintainability.
LLM evaluation is critical for generative AI in the enterprise, but measuring how well an LLM answers questions or performs tasks is difficult. Thus, LLM evaluations must go beyond standard measures of “correctness” to include a more nuanced and granular view of quality.
Key Topics:
3. Making Adoption Decisions:
Weighing cost, team familiarity, ecosystem integration, and performance.
Potential fallback or layering strategy (e.g., default to one assistant, then cross-check with another)
TFiR is one of the largest B2B YouTube channels focussed on emerging technologies.
Choosing the right generative AI tools is crucial for success. This course will teach you how to identify and evaluate leading platforms like OpenAI, Claude, and Gemini, considering key factors such as cost, privacy, scalability, and compatibility.
Read about layering multiple AI tools—like using one for generation and another for sanity checks.
Activity:
Implement a test harness (e.g., a series of automated unit tests) for your microservice.
Let your code assistant generate or improve some of these tests.
Run the tests across the code that was generated/refined by the AI in previous lessons.
Record how many tests pass/fail and whether the AI code needed extra fixes.
Goal:
Obtain a quantifiable measure of how well the AI’s code performs in your environment. This sets the stage for real-world adoption decisions.
Set Up Your Testing Framework
Decide which test framework you’ll use (e.g., Jest for Node.js, Mocha, Jasmine, or a Python framework like pytest).
Example (Node.js + Jest):
bash
CopiarEditar
npm install --save-dev jest
Add a test script to your package.json:
json
CopiarEditar
{
"scripts": {
"test": "jest"
}
}
Identify the Code to Test
If you followed the previous lessons, you have a microservice or endpoint (e.g., an Express server) with features the AI helped you build.
For Node.js, your functionality might be in app.js or a separate file with exported functions (e.g., userLogin.js, userRoutes.js).
Create a Test File
In your project, create a __tests__ folder or a file named something like app.test.js (Jest convention).
If you have a function, say userLogin, or an Express route, you can test them like this:
javascript
CopiarEditar
// app.test.js
const request = require('supertest');
const app = require('../app'); // or wherever your server is
describe('User Login Tests', () => {
it('should return success for correct credentials', async () => {
// ...
});
});
Ask Your Code Assistant to Generate/Improve Tests
Add a comment or partial code outline in app.test.js to see if your AI tool will auto-suggest test cases:
javascript
CopiarEditar
// I'd like a test for invalid credentials, missing fields, and success scenarios.
// Possibly test for error codes as well.
Accept the suggestions if they align with your logic. If not, refine them with more instructions:
javascript
CopiarEditar
// We have userLogin() that checks username/password
// Return "Login successful" or "Invalid credentials"
// Please generate tests for these scenarios:
The AI might propose multiple test blocks, e.g.:
javascript
CopiarEditar
describe('userLogin function', () => {
it('returns "Login successful" for correct creds', () => { ... });
it('returns "Invalid credentials" for wrong password', () => { ... });
it('handles missing username or password properly', () => { ... });
});
Run the Tests
In your terminal:
bash
CopiarEditar
npm test
Observe how many tests pass or fail.
If some fail, see whether the problem is in your AI-generated code or the tests themselves.
Fix any logic issues or test assumptions.
Document Pass/Fail Results
Keep a small record, e.g. TEST_RESULTS.md:
markdown
CopiarEditar
# Test Results
- `userLogin` tests:
- 3 tests passed, 1 failed initially (the AI had a mismatch in expected error message).
- After minor fix, all 4 pass.
- `userRoutes` integration:
- 2 tests pass, 1 fails due to missing field check. AI suggested partial validation.
- Manually added a `if (!req.body.username) ...` => all pass now.
Summarize if you needed to do “extra fixes,” or if the AI code worked mostly out-of-the-box.
Optional: Test the Entire Microservice
If you have an Express server, you can do integration tests using supertest or a similar library:
javascript
CopiarEditar
describe('GET /users', () => {
it('should return a list of users', async () => {
const res = await request(app).get('/users');
expect(res.status).toBe(200);
expect(res.body).toEqual(expect.any(Array));
});
});
Let the AI propose more advanced integration tests (e.g., test with JWT auth or different user roles).
Wrap Up
You’ve successfully integrated an automated test harness, generated or refined tests with your AI assistant, and executed them against AI-generated code.
You now have quantifiable pass/fail data and insight into how robust the AI’s code is in your environment.
Design a mini-benchmark using the microservice you’ve built so far:
Perform 2–3 typical tasks (like “add a /reports endpoint,” “refactor user auth,” “optimize database queries”).
Time yourself (or measure lines-of-code, commits) with and without AI assistance.
In a BENCHMARK_REPORT.md file, detail:
Task descriptions.
Time to completion (with AI vs. manual coding).
Test coverage or success rates for the AI-generated sections.
A recommendation on which assistant(s) you’d adopt and in what scenarios.
Objective: Learn how to run a straightforward but meaningful evaluation of different code assistants or approaches, then make an informed decision about which tool is most suitable for specific tasks or the entire project.
Choose 2–3 Typical Tasks
Decide on tasks that mirror real-world usage of your microservice. Examples:
Add a /reports endpoint returning aggregated data.
Refactor user authentication to add tokens or improve hashing.
Optimize an existing query (e.g., if you have a DB call that needs indexes or batch processing).
Ensure these tasks are big enough to require some coding, yet not so large that you can’t finish each in a short session.
Define Your Benchmark Approach
For each task, you’ll do it once manually (without AI) and once with AI (or vice versa).
Alternatively, if you have two AI assistants, you could do:
Task A with manual coding vs. AI
Task B with Assistant 1 vs. Assistant 2
Task C with a different approach or a mix.
Keep it consistent so that you can fairly compare the time and test results across tasks.
Prepare a Minimal “Stopwatch” or Time-Logging Method
You can use a manual timer on your phone or a simple script that logs start/stop times.
For each task, note:
Start coding time (minute:second).
End coding time when the feature is done or tests pass.
Perform Task #1 Without AI Assistance
Implement or refactor the code entirely by yourself (no code assistant suggestions).
Record:
Time to completion
Lines of code (if relevant)
Tests pass/fail or coverage changes
Perform Task #1 With AI Assistance
Reset your code or keep a new branch.
Implement the same feature with your code assistant.
Time how long it takes, track how many lines the AI helps with, note test coverage or pass rates.
Repeat for Task #2 and (Optionally) Task #3
If you have 2 or 3 tasks, do the same procedure: manual vs. AI, or AI vs. AI.
Stay consistent in how you measure time and coverage.
Document Everything in BENCHMARK_REPORT.md
Example Structure:
markdown
CopiarEditar
# Benchmark Report
## Task Descriptions
1. **Add /reports endpoint** to return user stats.
2. **Refactor user auth** to add JWT.
3. **Optimize DB queries** on `/analytics`.
## Time to Completion
| Task | Approach | Time (minutes) | LOC Written | Tests Passing (%) |
|-------------------|------------- |--------------- |------------ |-------------------|
| /reports endpoint | Manual | 15 | 50 | 95% |
| /reports endpoint | AI-Assisted | 10 | 40 | 95% |
| Refactor user auth| Assistant 1 | 12 | 30 | 90% |
| Refactor user auth| Assistant 2 | 14 | 35 | 85% |
...
## Observations
- **Task #1**: AI saved 5 minutes but required more code style fixes.
- **Task #2**: Assistant 1 had more robust security suggestions, while Assistant 2 was faster but missed some edge cases.
## Coverage or Success Rates
- For the AI-generated code, coverage was slightly lower until manual fixes.
- For manual coding, coverage remained stable but took extra time to write tests.
## Recommendation
- For rapid prototyping: use Assistant 1 if code style is less critical.
- For security-sensitive features: prefer Assistant 2 or manual approach with thorough reviews.
- Potential fallback: if coverage dips below 80%, do a manual pass or ask the AI to auto-generate tests.
Conclude with a Recommendation
Summarize which assistant(s) or approach might be best for certain tasks:
Speed vs. thoroughness
Simple features vs. complex refactors
Provide next steps: e.g., “We plan to adopt Assistant A for new microservices but keep manual checks for security-critical code.”
Share or Push Your Codebase
Ensure your final code for each approach is in your repository.
Upload BENCHMARK_REPORT.md alongside it:
bash
CopiarEditar
git add .
git commit -m "Add AI benchmark with 3 tasks and final recommendations"
git push
Let others review or provide feedback.