Skip to content
View MarcinMikula's full-sized avatar
🎯
Focusing
🎯
Focusing
  • Warsaw

Block or report MarcinMikula

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
MarcinMikula/README.md

Quality Assurance Engineer exploring AI-assisted quality engineering β€” 13+ years across telco, banking, insurance & e-commerce.

I use these repositories primarily as a learning and engineering workspace.

My goal is not to collect technologies or build impressive demos as quickly as possible. I use the projects to organise and deepen what I already know about software testing and automation, while deliberately expanding into Python, AI and LLM-based systems.

A large part of that learning happens through building and questioning things:

  • revisiting software-testing and automation fundamentals in working code,
  • improving my Python and engineering skills through real implementation problems,
  • learning how modern AI systems work by testing where they help, where they fail, and where deterministic controls are still necessary,
  • turning useful experiments into practical QA tools.

The result is gradually becoming a small, internally consistent testing ecosystem rather than a collection of unrelated repositories.

enterprise testing experience
        ↓
testing & automation fundamentals
        ↓
Python / Playwright / pytest / APIs
        ↓
LLM-assisted experimentation
        ↓
practical QA tools
        ↓
more questions β†’ more learning

πŸ”­ Selected engineering projects

Project What it does
πŸ”₯ PhoenixQA Experimental self-healing automation framework exploring locator recovery, Playwright actionability failures, LLM-assisted diagnosis and deterministic evidence-based guardrails.
πŸ› οΈ TestRepairEngine Focused installable runtime-recovery component for Playwright tests. Uses deterministic recovery first, bounded local-LLM escalation only when justified, and leaves the unchanged original test as the final oracle.
πŸ—ΊοΈ TestCartographer Human-guided context acquisition, automation adaptation and maintenance tool that maps application evidence into reviewed framework changes and reusable project knowledge.
πŸ—οΈ qa-automation-framework Reusable Python + pytest + Playwright automation skeleton built around Page Object and Service Object patterns for UI and API testing.

These projects explore different parts of the same automation lifecycle:

application knowledge
        ↓
TestCartographer
        ↓
qa-automation-framework
        ↓
normal test execution
        ↓
TestRepairEngine
        ↓
runtime recovery evidence
        └──────────────→ TestCartographer


PhoenixQA
        ↓
experimental research into broader
self-healing and actionability recovery

They have different responsibilities, but share the same underlying principle:

AI may propose and assist; evidence, deterministic controls and human decisions define what may be trusted.


πŸ§ͺ Research & supporting projects

  • llm-qa-toolkit β€” research prototype exploring Test Basis, gradability, evaluator authority and the limits of LLM-as-a-judge evaluation.
  • defect-pilot β€” experiment in Jira defect quality, enrichment and explicit completeness gates; earlier automated-retest ideas were deliberately narrowed after practical evaluation.
  • test-design-gatekeeper β€” evidence-grounded, ISTQB-informed functional test-case review assistant progressing through solution and architecture design.

🧠 How I work

I use AI heavily, but I don't treat generated code or generated conclusions as a black box.

My background is software testing rather than software development, so I approach these projects primarily through SDLC/STLC, risk, testability and evidence. I bring practical experience with enterprise systems, SQL, REST/SOAP integrations and test design, while deliberately expanding the programming side through Python, Playwright, pytest and related tooling.

For automation I try to preserve familiar engineering boundaries such as Page Objects, Service Objects, separation of test data and logic, deterministic checks, explicit acceptance criteria and regression evidence.

I rarely treat the first technically working solution as the final one. Before claiming that an approach works, I try to understand the alternatives, assumptions and failure modes behind it β€” and, where possible, test them against evidence.

That sometimes makes the projects slower and more exploratory than a typical implementation project. A feature may lead to a diagnostic experiment, an architectural change, a rejected approach or even a reduction in project scope. I consider that part of the engineering work rather than wasted effort.

I prefer a smaller claim supported by evidence over a more impressive claim supported only by working code.

A seemingly small implementation question can therefore turn into:

implementation
    ↓
unexpected behaviour
    ↓
test
    ↓
architectural question
    ↓
alternative approaches
    ↓
experiment
    ↓
documented limitation
    ↓
new hypothesis

That reasoning is deliberately preserved in files such as LEARNINGS.md, testing strategies, architecture decisions, known limitations and research hypotheses.

The repositories therefore show not only the current code, but also how and why it evolved β€” including approaches that were tested and later rejected or deliberately narrowed.


🌱 What I'm currently learning

  • practical limits and useful applications of LLMs in software testing,
  • reliable evaluation of AI systems and the weaknesses of LLM-as-a-judge,
  • self-healing and maintenance of automated tests,
  • human-controlled vs autonomous AI-assisted workflows,
  • Python and software-design practices needed to turn QA experiments into maintainable tools,
  • CI/CD, reproducibility and measurable validation of AI-assisted behaviour,
  • local vs cloud LLM trade-offs for enterprise and sensitive test contexts.

A recurring principle across the projects is:

LLM proposes; evidence, deterministic controls and humans define what may be trusted.


πŸ› οΈ Current toolbox

Python Playwright pytest SQL REST SOAP Postman Newman POM SOM Jira API Ollama Claude / Anthropic API Git GitHub Actions Allure


πŸ“ Based in Warsaw β€” open to remote/hybrid roles

nofluffjobs profile

Pinned Loading

  1. qa-automation-framework qa-automation-framework Public

    Reusable Python test automation framework skeleton for UI and API testing | Playwright Β· pytest Β· POM Β· SOM

    Python

  2. llm-qa-toolkit llm-qa-toolkit Public

    A framework for testing LLM-based chatbots in regulated industries (telco, banking, insurance). Covers hallucination detection, prompt injection resistance, response quality scoring and regression …

    Python 1 1

  3. defect-pilot defect-pilot Public

    Software quality starts with test process quality. Privacy-first AI QA gatekeeper for Jira and the system under test β€” enriches bug reports and enforces completeness standards.

    Python

  4. PhoenixQA PhoenixQA Public

    πŸ”₯ Self-healing test automation framework. When a Playwright selector breaks, PhoenixQA diagnoses the failure with LLM, proposes a fix, and learns from every decision. Local (Ollama) or API (Anthrop…

    Python