What does it mean to “find the right team”?
The project began with a deceptively simple challenge: explain TeamCreator to a technically serious person who keeps asking for the formula.
The question is fair. If a system makes a recommendation, we should be able to say what it takes as input, what it returns, how it can be wrong, and what evidence would change our mind.
But the request for one deterministic formula smuggles in an assumption that does not fit the object. A project team is not a closed physical system. The people, task, dependencies, tools, manager, time pressure, and prior working relationships change one another. A team’s previous coordination can become an input to its next phase. A missed handoff can create new workload. A new person can change the communication graph. The outcome is not simply the output of a static profile.
The first hypothesis seed came from a conversation at the boundary between product language and hard-science expectations. The proposed vocabulary was multivariate analysis, sensitivity analysis, non-linear interaction, Bayesian networks, ensembles, PCA, and what-if simulation. Those are useful families of methods. They are not evidence by themselves. A model name cannot rescue an undefined construct, a leaking label, or a dataset that only records successful projects.
“TeamCreator does not calculate the objectively correct team. It makes a particular team configuration inspectable for a particular project, at a particular time.”
It first checks feasibility. It then compares explainable scenarios, exposes evidence and unknowns, records the owner’s decision, and learns cautiously from later observations. The claim is about decision quality and scientific testability before it is about prediction.
From a universal formula to a conditional question
Instead of asking, “What is the score of this person?” we ask: “Given these project requirements, this time window, this evidence snapshot, this capacity state, and this collaboration context, which scenarios are feasible and what risks distinguish them?” The target is no longer an essence. It is a decision, attached to a time horizon and an outcome definition.
The supplied Atlantykron text therefore functions as a hypothesis generator. It motivates sensitivity analysis and multi-variable observation. It does not establish that communication frequency is quality, that response latency is commitment, or that a graph model will generalize across organizations. The research program is the work of turning attractive intuitions into operational definitions that can fail.
Established research
Peer-reviewed findings and authoritative methods that constrain what a responsible model may claim.
Product design
Deliberate boundaries such as hard constraints first, evidence provenance, owner approval, and visible uncertainty.
Pilot hypothesis
Testable predictions about feasibility, coordination, decision time, calibration, and correction.
Unknown / prohibited claim
Anything we cannot currently validate: universal success prediction, intrinsic person scores, or hidden trait inference.
Make every claim smaller than the evidence.
The discipline is not to sound scientific. It is to make the product’s statements falsifiable, time-bound, and traceable to what was actually observed.
The scientific method for TeamCreator starts before model training. We define a research question, specify the unit of analysis, translate vague concepts into measurable variables, freeze a data snapshot, choose a baseline, predefine outcomes, and record what changed. The interface is part of the study because the intervention is not “AI” in the abstract; it is a workflow that changes how an owner sees evidence and alternatives.
The TeamCreator method loop
A result is not “the model was accurate.” A result is “under this split, for this defined outcome and horizon, this method added this much information over this baseline, with these calibration and fairness limits.”
What “scientific” means in the product
Observable. “Coordination readiness” must resolve into variables such as named handoff ownership, prior shared work on the relevant interface, blocker age, dependency closure, or a voluntary team-process pulse. A label that cannot be collected consistently cannot be treated as a clean target.
Contextual. A pattern can be real in one organization and fail in another. Task interdependence, novelty, work mode, governance, manager practice, and resource constraints belong in the context layer, even if that makes the model less glamorous.
Time-indexed. Every forecast needs a horizon: the next sprint, the next project phase, the first 30 days, or another declared window. “Team success” without a time boundary is not a label.
Reproducible. The same evidence snapshot and rule/model version should replay to the same scenario explanation. If a result cannot be reproduced, we cannot tell whether the system learned or merely changed.
Correctable. Human corrections and overrides are signals about fit, evidence quality, and workflow. They are not automatically ground truth. The system must preserve the reason for a correction and the alternative path that was chosen.
The five research questions
Can structured requirements and capacity reveal fragile or infeasible teams earlier?
Compare hard-constraint and provenance-aware scenarios with skill-only checklists.
Do observable working interfaces add information beyond capability coverage?
Test prior shared episodes, handoff structure, and expertise discoverability under different interdependence levels.
Do team processes change across project episodes?
Use time-ordered team-level observations rather than permanent member labels.
Does evidence-linked scenario comparison improve owner decisions?
Measure time to a feasible team, correction loops, unowned gaps, and appropriate trust.
Where does calibration degrade, and who bears the data burden?
Hold out projects and contexts; audit missingness, proxies, opportunity, and safe abstention.
The product is built on team science, not personality mythology.
The literature gives us useful constructs, but each one has to be translated into an ethical, observable product mechanism.
Dynamic, multilevel teams
Kozlowski and Ilgen’s review, and Ilgen et al.’s IMOI model, treat teams as adaptive systems in which inputs, processes, emergent states, and outcomes interact across levels.
Temporal team processes
Marks, Mathieu, and Zaccaro distinguish transition, action, and interpersonal processes. They are useful for deciding what to instrument at what phase.
Psychological safety
Edmondson connects psychological safety with learning behavior in work teams. This supports asking whether bad news and help-seeking can surface.
Transactive memory
Research on transactive memory explains how teams know who knows what, and how expertise discoverability can support coordination.
Composition is contextual
Team-composition reviews and meta-analysis show that attributes and configuration interact with task context and interdependence.
Formation is optimization
Operations-research work formalizes coverage, communication cost, trust, geography, temporal collaboration, or proficiency as constrained objectives.
What these foundations do not permit
The literature does not justify reading character from a single message, inferring psychological safety from response time, replacing context with a personality embedding, or ranking people as universally good or bad. The responsible translation is deliberately less dramatic: instrument team processes, preserve context, show uncertainty, and create an opportunity for correction.
“Complex adaptive system” is therefore a useful explanatory lens, not a license for unfalsifiable language. For every feature we need an observable definition, collection method, time window, unit of analysis, missingness explanation, privacy boundary, and a declared relationship to a decision or outcome.
A readiness vector, not a magic number.
The first release should make feasibility and trade-offs visible before it tries to learn a universal score.
Let a project brief at time t be requirements, skills, proficiency or evidence thresholds, the time window and demand, work-mode constraints, and known risks. Let a person’s capability state include claims, levels, evidence, provenance, uncertainty, recency, consent, and visibility. Let a candidate team be a time-indexed collaboration graph whose edges mean observable working relationships—not liking, similarity, or inferred personality.
Ci(t) = (Ai, Li, Ei, Pi, Ui, Zi)
GT(t) = (VT, ET(t), ωT(t))
The notation forces the model to carry time, evidence, uncertainty, and consent alongside skills. A forecast is conditional on a context and horizon, never an intrinsic person value.
For a candidate team T, the output is a decomposable readiness vector:
Team readiness / example UI representation
These are explanatory components, not validated probabilities. A UI may sort scenarios with a weighted objective, but it must show which inputs and trade-offs produced the ordering.
Hard constraints first
If the project requires capability coverage, hours, a time window, role ownership, geography, conflict-of-interest checks, or consent, those are constraints. A weighted score must not turn an infeasible scenario into a reassuring number.
Σi xi(hi,p,t + δi,p,t) ≤ Hp,t
If no candidate satisfies the hard constraints, return “no feasible scenario under current evidence and constraints” and show the smallest documented relaxation that would make one possible.
Scenario comparison, not false precision
Once constraints pass, the system can compare objectives such as coverage, capacity safety, complementarity, evidence quality, coordination cost, dependency concentration, and uncertainty. The weights are scenario configuration choices, not discovered facts about human worth.
Coverage first
Maximizes demonstrated role and capability coverage. Useful when missing a critical skill is the dominant risk.
Capacity safe
Protects the delivery window and current commitments. Useful when over-allocation is the dominant risk.
Interface light
Reduces handoff and single-point-of-failure risk. Useful when work is tightly coupled or novel.
A Pareto set or a small family of distinct scenarios is more honest than a single ranking. The owner can see what is gained and sacrificed, record an override, and schedule a review when the unknown is resolved.
Fourteen hypotheses that can be wrong.
The hypothesis register is the center of gravity: every ambitious product claim is converted into a test, a baseline, and a failure condition.
We group the hypotheses into five study families. The early pilot should not attempt to estimate everything. It should establish whether the measurement layer is usable, whether the deterministic baseline is replayable, and whether owners can correct the system without hidden pressure.
Hard constraints improve feasibility
Constraint-aware, provenance-linked scenarios should produce fewer infeasible proposals than skill-only matching.
Capacity adds information beyond capability
Availability, committed hours, coordination load, and uncertainty should improve near-term feasibility decisions beyond capability coverage alone.
Prior shared work and discoverability improve coordination readiness
Working ties, named handoff ownership, and knowledge discoverability should matter more when task interdependence is high.
Complementarity beats raw diversity counts
Coverage of task interfaces and complementary capabilities should explain defined deliverables better than counting skills or demographic categories.
Team processes change across episodes
Transition, action, and interpersonal indicators should vary by project phase and predict the next operational window without becoming person labels.
Recency and confirmation change uncertainty
Evidence age, directness, source, and human confirmation should change confidence in a claim more reliably than they change “ability.”
Decision Room improves decision quality and speed
An evidence-linked workflow should reduce time to a feasible team, clarification loops, and unowned gaps while preserving owner authority.
Pareto scenarios beat one weighted score
Showing distinct trade-offs should produce better-informed owner choices than one opaque or over-compressed ranking.
Calibration degrades across contexts
A model or heuristic calibrated on one project type, organization, or collaboration regime should degrade in another.
Missingness is informative and unequal
Missing evidence can mean no profile, no consent, connector failure, new work, or irrelevance; treating it as negative creates systematic errors.
Hybrid explainability has practical value
A rule-plus-statistical model may match or exceed opaque model utility while improving correction, trust calibration, and reviewability.
Context explains actionable variance
Team and project context should explain more actionable variation than stable individual traits.
Corrections are learning signals, not ground truth
Human edits reveal evidence gaps, workflow mismatches, and preference differences; they need reasons and outcome follow-up.
Safe abstention improves decisions under uncertainty
Declining to rank or recommending evidence collection should outperform confident guesses when critical inputs are missing or out of distribution.
The full hypothesis register is maintained in the project research dossier. The page presents the public narrative; the source register retains operational definitions, designs, and analysis notes.
Data is not a pile. It is a contract.
A sophisticated model trained on ambiguous, stale, or post-outcome data is less scientific than a transparent baseline with clean provenance.
The data contract separates raw inputs, extracted claims, approved facts, decisions, and later outcomes. That separation lets us answer the questions a reviewer will ask: what did the system know at the time, who confirmed it, what was missing, what did the owner choose, and when did the outcome occur?
Public datasets: useful scaffolding, not proof
The AMI Meeting Corpus can support work on meeting interaction and multimodal coordination. SocioPatterns and MIT Reality Mining provide examples of proximity or interaction-network data. GHTorrent and PROMISE can support software-engineering process and outcome experiments. They cannot validate TeamCreator’s organizational use case: their populations, tasks, consent boundaries, labels, and collection regimes differ. They are method supplements, not a substitute for a consented internal pilot.
The model ladder
Hard constraints, transparent score components, evidence links, missingness and abstention.
Regularized regression, mixed effects, calibration, survival or count models with time and project clustering.
Time windows, handoff networks, dependency centrality, sequence features, and drift monitoring.
Gradient boosting, link prediction, graph embeddings, or GNNs as research candidates and candidate generators.
Only when the estimand, assignment unit, time horizon, contamination, and missing outcomes are specified.
Why not start with deep learning? Because model complexity is not an answer to scarce, clustered, biased, or post-selection data. A larger model can memorize organization identity, encode tool-access proxies, and produce a persuasive explanation after the fact. The research program earns complexity by showing incremental utility over the deterministic and interpretable baselines.
The baseline is part of the result.
A model is promoted only when it improves validated outcomes, calibration, explanation, and safety over what came before it.
Offline checks
- Hard constraints pass or fail explicitly; no unsupported claim is silently counted as coverage.
- Every displayed component links to an evidence fragment or a named rule.
- The same snapshot and version replay to the same result.
- Infeasible briefs generate a useful gap explanation and smallest documented relaxation.
- Learned models use project/time splits, not random rows, and keep all records from a project in one fold.
- Probabilistic outputs report calibration, Brier score or log loss where appropriate, uncertainty, abstention, and out-of-distribution flags.
- Ranked scenarios report constraint violations, top-k utility, owner correction, diversity of alternatives, and hindsight regret—not only NDCG or precision.
Prospective human-in-the-loop evaluation
The meaningful first study compares the existing decision process with the TeamCreator workflow without removing the owner’s authority. Candidate measures are time from brief to first feasible scenario, clarification loops, infeasible proposals caught before commitment, evidence coverage, correction and override rates, unowned risk count, next-step completion, later 30-day and 90-day observations, explanation clarity, appropriate trust, and whether people can correct or withdraw their evidence.
Minimum pilot protocol
Project brief or project-phase decision, with a declared owner and time window.
Existing workflow or skill-only checklist versus deterministic, evidence-linked scenarios.
Decision Room workflow: structured brief, evidence links, gaps, alternatives, owner rationale, review date.
Feasibility, decision time, correction quality, coordination signals, and defined delivery observations.
Predeclare split, estimand, missing outcome handling, stopping rule, and context strata.
Human approval, visible uncertainty, no adverse employment decisions, withdrawal and correction path.
Causal caution
If TeamCreator helps choose teams, selected teams are not random. Better projects may go to stronger managers, receive more resources, or be easier to observe. A correlation between a team feature and delivery is not automatically a causal effect. A future intervention study must predefine the estimand, unit of assignment, outcome window, contamination risk, and treatment of missing outcomes.
Responsible use is an engineering requirement.
The model is not safe because the copy sounds cautious. The data path, UI, audit log, and human alternative have to enforce the boundary.
What TeamCreator must not become
It is not an automated hiring, firing, promotion, disciplinary, or allocation decision. It is not a psychometric system. It does not infer health, disability, emotion, nationality, religion, political preference, or personality from private or incidental signals. It does not treat message sentiment, facial expression, meeting attendance, code volume, or response latency as universal measures of commitment or worth.
The governance contract includes provenance, timestamp, extraction method, confidence, confirmation state, correction, retraction, deletion, export, consent withdrawal, purpose limitation, a human alternative when AI is unavailable or contested, and an audit record for requirement, evidence, rule/model, decision, and override versions. Legal review remains necessary for the actual jurisdiction and deployment context; standards are inputs to governance, not a substitute for counsel.
If a critical fact is unknown, stale, conflicting, or outside the model’s context, the safest output may be a question—not a ranking.
“Unknown” is not a failure of the person. It is a property of the evidence state. The system should say what is missing, why it matters, who can resolve it, and what decision is blocked until then.
Start with a small study that can survive scrutiny.
The next safe artifact is a replayable Brief → Evaluator → Composer → Decision Record loop, not a claim that machine learning has solved team formation.
Before the Atlantykron presentation or any model training, we need three to five golden project briefs, six to ten consented synthetic or real pilot profiles, a deterministic feasibility engine, and a record of every evidence fragment, correction, decision, unknown, and later observation. The first achievement is measurement quality: can two reviewers understand the same scenario and disagree for a visible reason?
What a scientific advisor can challenge
The right invitation is not “please endorse our AI.” It is “please attack the assumptions in our measurement and validation pipeline.” A methodological advisor can help decide how to handle collinearity and spurious correlations, how to separate signal from workplace noise, how to specify boundary conditions for sensitivity analysis, whether a Bayesian network is warranted, and what evidence would invalidate the product’s current framing.
“We are building a multivariate, evidence-backed sensitivity model for team-project decisions. We want an expert to critique the data assumptions, feature definitions, validation splits, and model-promotion gates before we claim predictive value.”
The proposed collaboration can be a 30-minute methodological deep dive, followed by a small advisory review of the data contract and first pilot protocol. The advisor does not need to bless a fixed model; the value is in choosing and rejecting architectures with evidence.
The Atlantykron explanation
TeamCreator asks whether a particular team configuration is feasible for a particular project at a particular time, with what evidence and what risks. It enforces hard constraints, compares several transparent scenarios, lets a human owner choose or modify one, records the decision, and observes what happens later. A Bayesian network, ensemble, temporal model, or graph neural network may eventually be tested. None is a justified production commitment until repeated, consented, time-indexed data shows incremental utility over a simpler baseline.
We are not building a formula for human teams. We are building the measurement system that lets an organization discover which small changes in requirements, capacity, interfaces, or ownership are associated with meaningful changes in a defined delivery window—and lets the organization see when the evidence is too weak to say.
Build the evidence loop
Brief, claims, evidence fragments, capacity, scenarios, decisions, corrections, and outcomes with provenance.
Compare against skill-only and naive baselines
Replay golden briefs; report feasibility, constraint violations, explanation coverage, and scenario diversity.
Run a prospective, human-in-the-loop pilot
Measure time, corrections, gaps, appropriate trust, and later team-level process outcomes.
Earn temporal, network, and causal complexity
Promote models only when context, calibration, uncertainty, governance, and incremental value are demonstrated.
The record stays open to correction.
This public page is the narrative layer of the research package. The bibliography below is the chain of ideas we are translating, not a claim that every cited result transfers directly to TeamCreator.
1,000 article and review records, collected from 24 overlapping query families.
The searchable corpus exposes the abstract, provenance, automated triage label, evidence-design signals, and likely TeamCreator pathway for every selected record. It is a discovery layer and a prepared human-review queue—not a claim that 1,000 full papers have been independently adjudicated.
Explore the 1,000-abstract corpus →RO / RU: fit before management, formal doctoral prospectus, technical method ladder, networking scripts, and evidence-bounded research lineage.
Use it to explain why TeamCreator exists, what we can claim today, and what a serious multi-study research program would need to prove next.
Open the RO / RU research companion →References [1]–[10] establish the team-science foundations: dynamic and multilevel teams, IMOI, psychological safety and learning, transactive memory, team effectiveness, temporal processes, composition, interdependence, and the level of analysis. References [11]–[14] establish the operations-research and network-formation families that motivate hard constraints, communication cost, trust, proficiency, and multi-objective scenarios. References [15]–[17] are governance and employment-impact inputs. References [18]–[22] are methodological dataset and repository anchors. The larger literature register is available through the corpus companion and working-repository artifacts.
[1] Kozlowski & Ilgen (2006), “Enhancing the effectiveness of work groups and teams,” Psychological Science in the Public Interest.
[2] Ilgen, Hollenbeck, Johnson & Jundt (2005), “Teams in organizations: From input-process-output models to IMOI models,” Annual Review / author PDF.
[3] Edmondson (1999), “Psychological safety and learning behavior in work teams,” Administrative Science Quarterly.
[4] Austin (2003), “Transactive memory in organizational groups,” Journal of Applied Psychology.
[5] He & Hu (2021), “Shared leadership and team creativity,” Journal of Business Research.
[6] Mathieu, Maynard, Rapp & Gilson (2008), “Team effectiveness 1997–2007,” Journal of Management.
[7] Marks, Mathieu & Zaccaro (2001), “A temporally based framework and taxonomy of team processes,” Academy of Management Review.
[8] National Academies (2015), “The science of team science,” Enhancing the Effectiveness of Team Science.
[9] Bell, Brown, Colaneri & Outland (2018), “Team composition and the ABCs of teamwork,” American Psychologist.
[10] Gully, Incalcaterra, Joshi & Beaubien (2002), “A meta-analysis of team-efficacy, potency, and performance,” Journal of Applied Psychology.
[11] Berktas & Yaman (2021), “A branch-and-bound algorithm for team formation on social networks,” INFORMS Journal on Computing.
[12] “A unified framework for effective team formation in social networks” (2021), Expert Systems with Applications.
[13] “Towards realistic team formation in social networks based on densest subgraphs” (WWW 2013), ACM Digital Library.
[14] “Effective team formation in collaboration networks using vertex and proficiency similarity measures,” AI Communications.
[15] NIST (2023), AI Risk Management Framework.
[16] European Union (2024), Regulation (EU) 2024/1689.
[17] U.S. EEOC, guidance on adverse impact and AI in employment selection.
[18] AMI Consortium, AMI Meeting Corpus.
[19] SocioPatterns, dataset catalogue.
[20] MIT Human Dynamics Laboratory, Reality Mining.
[21] Gousios (2012), “The GHTorrent dataset and tool suite,” IEEE MSR.
[22] PROMISE Software Engineering Repository, official catalogue.
Research artifacts in the working repository
The narrative is backed by a dated scientific basis, a data-and-model contract, a hypothesis register, a BibTeX file, and standalone diagrams. They are maintained as source artifacts so the public story can change when the evidence changes.