The science

Structured couples work helps. Here’s how much.

Every number on this page comes from a peer-reviewed study we have read and checked. The limits are on the page too, including the paper that argues against our whole premise.

None of it is a claim about Thirdlight. We have no outcome data of our own yet, and we say so plainly below.

Does structured couples work actually help?

A 2020 meta-analysis of 58 studies covering 2,092 couples found a large improvement in relationship satisfaction from before to after treatment (Hedges’ ḡ = 1.12), which is a within-group gain and not the same thing as a controlled effect.

The short answer is yes. The honest version has a footnote attached to it.

Pooled across the field, couples who go through a structured program come out better than they went in, and the size of that change is large. The catch is in the measurement. A within-group number compares a couple to their own starting point. It does not compare them to couples who got nothing, or to couples who got something else, and those comparisons always come out smaller.

Make the harder comparison and the effect holds at a more modest size. A 2024 meta-analysis of emotion-focused couples work reports d = 0.44 against other active couples interventions, and d = 0.93 against nothing at all. We quote the first number.

A smaller study followed 32 couples for two years after treatment and found the gains kept improving slowly instead of fading. There was no control group at that stage, so read it as encouraging rather than settled.

What the study did

58 studies, 40 independent samples, 2,092 couples, pooled with random-effects modelling and moderator analyses across design and timeframe. It is the current standard reference for the question. The emotion-focused meta-analysis included randomized trials, quasi-experimental work and unpublished dissertations, which lowers publication bias and lets in some weaker studies at the same time.

What it doesn’t show

  • The headline figure is a within-group gain. Read it as “couples improved”, not as “the program caused the improvement”.
  • The long-term picture is more sober. 134 seriously distressed couples were randomly assigned to about eight months of in-person couples work and followed for five years. Gains were large, held for roughly two years, then eroded.
  • Almost none of this was measured on couples arriving the way ours do, at the dose we offer.

If you’re the one having the fight

  • If you are deciding whether structured help is worth the hours, the field’s answer is that it moves the thing you care about, by an amount you would notice.
  • It also says the useful measure is where you are months later, not how you feel walking out of one good conversation.

What we do with it

Gains fade without upkeep, so the weekly slot is a standing appointment rather than a course you finish.

Does it still work through a screen?

A 2025 meta-analysis of six randomized trials found digital couples interventions produced a moderate improvement in relationship satisfaction (Hedges’ g = 0.42), with substantial variation between the trials pooled (I² = 66%).

This is the question the partner who did not sign up asks first, and it is a fair one.

The best single answer is a trial run across the United States in 2020 with 742 relationship-distressed, low-income couples. Median household income was about $27,000. Couples were randomly assigned to one of two web-based programs or to a waitlist, and both programs improved relationship functioning. That is a group earlier relationship-education efforts had barely moved.

The pooled picture agrees with it. Fifteen randomized trials of technology-delivered couples interventions have now been reviewed, and the six with combinable data produce a moderate effect on satisfaction.

What the study did

742 couples, 1,484 individuals, recruited nationwide rather than through one clinic, randomly assigned, with a waitlist control. Both programs were mostly self-directed online activities worked through at home. Both also included short coaching calls with a human being.

What it doesn’t show

  • A person was in the loop. Both programs included coaching calls, so this is not evidence that software on its own works, and we would be overreading it if we said otherwise.
  • The comparison was a waitlist, which means getting nothing. Effects measured against nothing always come out larger than effects measured against a real alternative.
  • The 2025 pooled estimate rests on six studies with high heterogeneity, and its authors flag language bias and methodological problems in several of the trials included.

If you’re the one having the fight

  • The format is not the obstacle most people assume it is. Doing something structured at home has better evidence behind it than waiting until things are bad enough to justify an office.
  • The entry cost is lower than the version people imagine, which is usually a year of weekly appointments.

What we do with it

Elena is a live conversation rather than a worksheet, which is closer to the thing that was tested. Whether an AI stands in for the coaching call has not been tested, by us or by anyone.

Is it normal for a relationship to get worse?

Across 165 samples and 165,039 people, relationship satisfaction declines from about age 20, reaches its low point around age 40 and rises into later life, though an average curve like that says nothing about where any individual couple is headed.

This is the most useful finding on the page, and close to nobody in this category publishes it.

Satisfaction drops across the first decade of a relationship, levels off, and then turns back up. It happens to most couples. It is not a verdict on yours, and it does not mean either of you did something wrong.

It also says something about timing. Couples in their thirties, inside the first ten years, are on the descending part of the curve. That is not a reason to panic. It is a reason to treat the drift as ordinary and do something about it while it is still drift.

What the study did

A systematic review and meta-analysis of 165 independent samples and 165,039 participants, combining longitudinal and cross-sectional data. Methodologically careful, and one of the best-evidenced facts in this whole area.

What it doesn’t show

  • An aggregate curve describes populations. It cannot tell you whether your relationship is improving or declining, and nobody should use it to.
  • The samples are predominantly Western.
  • “It goes back up later” describes an average across a lifetime. It is not a reason to wait.

If you’re the one having the fight

  • If things feel worse than they did five years ago, you are on the same curve as most people. That is information, not a diagnosis.
  • The dip being normal is not the same as the dip being fine. Normal things still cost you years.

What we do with it

The written note after each session and the trajectory it builds exist so you are looking at your own line instead of the average one.

Why does one of us push while the other goes quiet?

A meta-analysis of 74 studies covering 14,255 people found a moderate, consistent association between the push-and-quiet pattern and worse individual and relational outcomes (r = .36), the same size whichever partner does the pushing.

Most couples who come to us describe the same argument twice. One person raises something. The other gets shorter, flatter, or leaves the room. The first pushes harder because the silence feels like being dismissed. The second shuts down further because the pushing feels like being cornered.

Researchers call this demand/withdraw. Elena calls it the push-and-quiet loop.

A 1990 study pinned down how it works. Thirty-one couples had two conversations on the same day: one about something he wanted her to change, one about something she wanted him to change. The roles moved with the topic. Whoever wanted the change did the pushing. Whoever was being asked to change did the retreating.

So the loop follows the topic. It is not a fact about who either of you is, and the pooled evidence since has held that line: both directions of the pattern carry the same weight.

What the study did

31 couples, videotaped in a lab, in the United States in 1990. Small, and the design is unusually strong for its size: each couple served as its own control across the two topics, and the conversations were rated by both partners and by observers who did not know what was being tested. A longitudinal companion followed 48 couples and found the pattern on the wife’s issue predicted lower satisfaction for wives a year later. The pooled association is stronger in distressed samples (r = .41) than in non-distressed ones (r = .35).

What it doesn’t show

  • The 1990 study is 31 couples in a lab. It shows the pattern clearly. It cannot tell you how common it is outside that sample.
  • Almost all of this work is correlational. Couples who do more of this are less happy. That is not the same as proving it is the engine.
  • It does not put the blame on the withdrawer, or on the pursuer. Some studies give the pushing side slightly more weight; the largest pooled estimate gives them the same.

If you’re the one having the fight

  • Neither of you started it, and each half keeps the other half going. That is why being the one who stops feels impossible.
  • Asking the quiet one to talk more rarely works. What works is changing the shape of the conversation, so the person being asked to change has a reason to stay in it.
  • You can usually predict which role you will take by asking who wants the change.

What we do with it

Elena names the loop out loud, both halves at once, attached to the topic rather than to a person. The raised issue gets a bounded turn, and the other person’s turn is scheduled before the first one starts.

Does it matter how a conversation starts?

A 1998 study of 130 newlywed couples followed for six years found that how an issue got raised in the opening minutes tracked with where the marriage went, from a model fitted to the same couples it was tested on, so read it as a description of what differs rather than a forecast.

There is a version of relationship advice that says the skill is listening. You repeat back what your partner said, you validate the feeling, then you take your turn. It is in most of the books.

That 1998 study tested the idea against several other candidates and it did not hold up. The couples who stayed happy almost never did anything resembling formal active listening while they were upset. It was mostly absent from the tapes.

What separated them was simpler. They opened issues gently, describing a situation and a feeling instead of leading with a verdict. When things got hot, someone made a small move that brought the temperature down and the other person let it work. Their bodies settled during the conversation. And the husbands in that sample took their wives’ positions seriously enough to give ground.

A 1999 study added the other half of it. Couples were recorded in a conflict conversation and then in a warm one straight afterwards. The couples who could put the fight down and re-enter warmth were on the better track. How a couple recovers carried information the fight itself did not.

What the study did

130 newlywed couples, one fifteen-minute conversation about a live disagreement, coded second by second for face, tone and content by people who knew nothing about them, then six years of follow-up. A good sample for observational work of its era. The classification percentages both papers report were fitted and tested on the same couples, so we do not quote them.

What it doesn’t show

  • Prediction figures in this literature do not survive being tested on new couples. A 2001 reanalysis showed accuracy drops substantially under proper cross-validation, and we think that critique is right.
  • The active-listening result is a failure to find something, which is weaker evidence than finding something. It does not prove structured listening is worthless.
  • The sample is 130 newlywed couples in the American Pacific Northwest in the early 1990s, mostly heterosexual and married. The influence finding was specifically about husbands in that sample.
  • The 1999 recovery study is small and its accuracy figures have the same problem. The durable part is the qualitative finding, not the percentage.

If you’re the one having the fight

  • The first three minutes are doing more work than the next thirty. If you open with “you always”, the conversation has largely happened already.
  • Rewriting your opening is small, specific and learnable, and it is probably the highest-leverage thing on this page.
  • Accepting influence is not giving in. It is finding the part of what your partner said that is fair and saying so out loud, before you make your own case.

What we do with it

Elena has you say the thing to her first, with a feeling and a request instead of a verdict, and the note carries the sentence you settled on into next week. She watches for repair attempts and stops to mark the ones that land. Our structured turn-taking is scaffolding that keeps two people from talking over each other, and we do not claim it is the mechanism.

Can something this small change anything?

In a randomized two-year study of 120 couples, three seven-minute writing exercises during the second year stopped a decline that kept going in the control group, an effect on the rate of decline rather than an improvement.

Everyone declined in year one. Nobody got anything that year, and that baseline is what makes the design convincing. The study watched both groups slide before it did a thing.

At the start of year two, half the couples were randomly assigned to a writing exercise. Three times over the year, seven minutes each, they wrote about a recent disagreement from the point of view of a neutral third party who wants the best for both of them. They also wrote about what makes that hard in the moment and how they would handle it. Twenty-one minutes across twelve months.

In year two the control group kept declining. The other group stopped.

The researchers traced the mechanism, and it is the part worth keeping. The writing did not make couples fight less often. It made the fights hurt less afterwards, and that is what held the line.

What the study did

120 couples, one site, two years, with a full year of pre-intervention baseline before random assignment. Small, and unusually persuasive for its size because of that baseline.

What it doesn’t show

  • The effect is on the rate of decline. The intervention couples did not report getting happier. They reported holding steady while the others slid.
  • One site, 120 couples, and brief psychological interventions have had a mixed replication record over the last decade. Promising, not settled.
  • The couples were doing reasonably well. This is a maintenance finding. There is nothing here about seriously distressed couples, and we would not offer it as a substitute for real help.
  • It was three spaced repetitions across a year. Doing it once and stopping is not what was tested.

If you’re the one having the fight

  • The dose that does something is far lower than the version people picture when they imagine getting help.
  • The goal worth aiming at is not fewer fights. It is fights that stop costing you three days afterwards.
  • There is nothing stopping you doing the exercise tonight. Take a recent disagreement, write about it as someone who cares about both of you and has no stake in who is right, then write about why that is hard in the moment.

What we do with it

The between-session check-in asks for one disagreement described from the outside. Seven minutes, spaced across weeks, rather than an hour in one go. The neutral third party the study asked couples to imagine is the seat Elena already occupies.

What happens when something goes right?

In a study of 79 dating couples with observer-coded conversations and a two-month follow-up, how a partner responded to good news was more strongly related to relationship well-being and to breaking up than how they responded to bad news, in a design that is correlational.

Almost everything written about couples is about trouble, including most of our own product. This is the finding that says the field is watching the wrong half of the day.

Seventy-nine dating couples took turns in a lab telling each other about a recent good thing and a recent bad thing. The person disclosing rated how understood and cared for they felt. Trained observers who knew nothing about the couples coded what the listener actually did. Two months later the researchers checked how the relationships were doing, including which had ended.

Responses to the good news mattered more.

Earlier work from the same group found the same shape. Telling someone about something good adds to your day beyond the good thing itself, and how much it adds depends on what the other person does with it. Curious, engaged responses produced the benefit. Flat ones, and ones that found the downside, did not.

What the study did

79 dating couples, videotaped, measured by both self-report and blind observer coding, with a two-month prospective outcome. Two independent measurement sources agreeing is the real strength here.

What it doesn’t show

  • 79 dating couples rather than married ones, over two months. Small and short.
  • It is correlational. Couples who respond well to good news are doing better. We cannot say from this that the responding causes the doing better.
  • It does not mean support during hard times is unimportant. The same study found responses to bad news mattered too. It found the good-news ones mattered more, in that sample, over that window.
  • It is not a script. Performed enthusiasm reads as performed.

If you’re the one having the fight

  • The moments carrying the most weight per second are not the fights. They are the twenty seconds after one of you says “you’ll never guess what happened”.
  • A flat “that’s nice” while looking at a phone is not neutral. Almost nobody does it on purpose. They are tired, or mid-task, or waiting for their turn to talk.
  • This is the cheap one. It takes less out of you than a hard conversation and it comes round several times a week.

What we do with it

The weekly session opens on one specific thing each of you was glad about, and Elena rewinds when good news goes past unmarked. We do not yet measure whether good news is landing in a couple’s week, which this research suggests is the more sensitive instrument. That one is planned.

Elena, as she appears in a session

From your coach

The research tells me where to look. It doesn’t tell me what I’ll find with the two of you.

Elena, your coach

Every pattern on this page shows up differently in a real conversation. Finding out which version you have is what the hour is for.

The limits

What the evidence doesn’t say

No service in this category publishes a caveat on its own evidence. Here is ours, and it is the part of the page we would most like you to read.

None of this is a result about Thirdlight.
We have no outcome data of our own. Nobody has completed a course of sessions with us, because there are no users yet. Every number above belongs to the category, and none of it transfers to us automatically. Ours has to be earned separately, and when we have it we will publish it the same way, including the parts that go badly.
The digital evidence includes a human being.
Both programs in the 742-couple trial had coaching calls with a person. Anyone citing that trial as proof that software alone works is overreading it, and so would we be. Elena is a live conversation rather than a worksheet, which is closer to the tested thing than a content library would be. Whether an AI substitutes for the coach on that call is untested.
Nobody can predict your relationship.
The divorce-prediction accuracy figures that circulate in this area come from models fitted to the same couples they were tested on. A 2001 reanalysis showed accuracy falls a long way under proper cross-validation. We do not quote those percentages, and it is reasonable to be suspicious of anyone who does.
Better communication may not be the cause.
The paper that argues against our whole premise is on this page on purpose. Newlywed couples were observed four times at nine-month intervals. More satisfied couples did communicate better at any given moment. But communication was a poor guide to how satisfaction changed over time, and the data fit “satisfaction drives communication” at least as well as the reverse. We teach conflict skills because skills are trainable and the process evidence is real. We cannot promise that better communication causes lasting satisfaction.
Gains erode.
The best long-term trial followed 134 seriously distressed couples for five years after about eight months of live, in-person work. Effects were large and largely held for two years, then satisfaction slid back. Nothing here holds a straight line forever, and anything promising otherwise is selling.
We work on one leg of a three-legged model.
The standard framework in this field says relationship quality comes out of the vulnerabilities each person brings, the stress they are under, and how they adapt together. We work on the third one. We cannot change what either of you arrived with, and we cannot remove your stressors.
The weekly session is built from research, not proven by it.
There is no controlled trial of a recurring structured couple meeting. We went looking. The weekly slot is built from work on shared activities that both partners genuinely want to be at, plus the small-repeated-exercise evidence above. That is what it is built from, and it is not the same as being proven by a trial.
Two claims we will not publish.
The 5:1 positivity ratio has no clean journal source. It is popularised in trade books rather than stated as a result in the paper it gets attributed to, so we cite the balance principle and never the number. And a widely circulated 2016 meta-analysis of web-based couples programs does not exist; what exists is a review article, which is a different kind of thing. Those are exactly the claims a competitor publishes and an assistant later corrects.

When to see a person instead

Thirdlight is a coaching product, not a replacement for therapy or medical care. If there is violence, coercive control, untreated addiction, or an affair that is still going on, none of this is the right tool. Elena will tell you that rather than keep you here. If you or someone you love is in danger, call or text 988 in the US, or your local emergency number.

The papers

Sources

Every paper cited above, in one list. Each was checked against the publisher record or PubMed while this page was written. If a claim here can’t be traced to one of these, it shouldn’t be here.

  1. Roddy, M. K., Walsh, L. M., Rothman, K., Hatch, S. G., & Doss, B. D. (2020). Meta-analysis of couple therapy: Effects across outcomes, designs, timeframes, and other moderators. Journal of Consulting and Clinical Psychology, 88(7), 583–596.

    doi.org/10.1037/ccp0000514Backs: Does structured couples work actually help? The pooled pre-post gain, and why we label it a within-group number.

  2. Spengler, P. M., Lee, N. A., Wiebe, S. A., & Wittenborn, A. K. (2024). A comprehensive meta-analysis on the efficacy of emotionally focused couple therapy. Couple and Family Psychology: Research and Practice, 13(2), 81–99.

    doi.org/10.1037/cfp0000233Backs: Does structured couples work actually help? d = 0.44 against active comparators, which is the number we quote.

  3. Wiebe, S. A., Johnson, S. M., Lafontaine, M.-F., Burgess Moser, M., Dalgleish, T. L., & Tasca, G. A. (2017). Two-year follow-up outcomes in emotionally focused couple therapy: An investigation of relationship satisfaction and attachment trajectories. Journal of Marital and Family Therapy, 43(2), 227–244.

    doi.org/10.1111/jmft.12206Backs: Does structured couples work actually help? 32 couples followed for two years, with gains continuing slowly.

  4. Christensen, A., Atkins, D. C., Baucom, B., & Yi, J. (2010). Marital status and satisfaction five years following a randomized clinical trial comparing traditional versus integrative behavioral couple therapy. Journal of Consulting and Clinical Psychology, 78(2), 225–235.

    pubmed.ncbi.nlm.nih.gov/20350033Backs: What the evidence doesn’t say The five-year erosion, and the ceiling any claim of ours sits under.

  5. Doss, B. D., Knopp, K., Roddy, M. K., Rothman, K., Hatch, S. G., & Rhoades, G. K. (2020). Online programs improve relationship functioning for distressed low-income couples: Results from a nationwide randomized controlled trial. Journal of Consulting and Clinical Psychology, 88(4), 283–294.

    doi.org/10.1037/ccp0000479Backs: Does it still work through a screen? The 742-couple trial, and the coaching calls inside it.

  6. Kernová, L., Halamová, J., & Deriglazov, D. (2025). Effectiveness of digital interventions on relationship satisfaction among couples: A systematic review and meta-analysis. BMC Psychology, 13, 1069.

    Open access · PMC12482273Backs: Does it still work through a screen? g = 0.42 for digital delivery, with the heterogeneity attached.

  7. Bühler, J. L., Krauss, S., & Orth, U. (2021). Development of relationship satisfaction across the life span: A systematic review and meta-analysis. Psychological Bulletin, 147(10), 1012–1053.

    doi.org/10.1037/bul0000342Backs: Is it normal for a relationship to get worse? 165,039 people, and the shape of the normal decline.

  8. Christensen, A., & Heavey, C. L. (1990). Gender and social structure in the demand/withdraw pattern of marital conflict. Journal of Personality and Social Psychology, 59(1), 73–81.

    doi.org/10.1037/0022-3514.59.1.73Backs: Why does one of us push while the other goes quiet? The two-topic design showing the roles follow the topic.

  9. Schrodt, P., Witt, P. L., & Shimkowski, J. R. (2014). A meta-analytical review of the demand/withdraw pattern of interaction and its associations with individual, relational, and communicative outcomes. Communication Monographs, 81(1), 28–58.

    doi.org/10.1080/03637751.2013.813632Backs: Why does one of us push while the other goes quiet? 74 studies, 14,255 people, and the equal weight of both directions.

  10. (1998). Predicting marital happiness and stability from newlywed interactions. Journal of Marriage and the Family, 60(1), 5–22.

    doi.org/10.2307/353438Backs: Does it matter how a conversation starts? Soft start-up, de-escalation and accepting influence.

  11. (1999). Rebound from marital conflict and divorce prediction. Family Process, 38(3), 287–292.

    doi.org/10.1111/j.1545-5300.1999.00287.xBacks: Does it matter how a conversation starts? Recovery carrying information the fight itself does not.

  12. Heyman, R. E., & Slep, A. M. S. (2001). The hazards of predicting divorce without crossvalidation. Journal of Marriage and Family, 63(2), 473–479.

    Open access · PMC1622921Backs: Does it matter how a conversation starts? Why we never quote a divorce-prediction accuracy figure.

  13. Finkel, E. J., Slotter, E. B., Luchies, L. B., Walton, G. M., & Gross, J. J. (2013). A brief intervention to promote conflict reappraisal preserves marital quality over time. Psychological Science, 24(8), 1595–1601.

    doi.org/10.1177/0956797612474938Backs: Can something this small change anything? Twenty-one minutes of writing, and the halted decline.

  14. Gable, S. L., Gonzaga, G. C., & Strachman, A. (2006). Will you be there for me when things go right? Supportive responses to positive event disclosures. Journal of Personality and Social Psychology, 91(5), 904–917.

    doi.org/10.1037/0022-3514.91.5.904Backs: What happens when something goes right? Responses to good news outweighing responses to bad news.

  15. Gable, S. L., Reis, H. T., Impett, E. A., & Asher, E. R. (2004). What do you do when things go right? The intrapersonal and interpersonal benefits of sharing positive events. Journal of Personality and Social Psychology, 87(2), 228–245.

    doi.org/10.1037/0022-3514.87.2.228Backs: What happens when something goes right? The listener’s response being where the value sits.

  16. Lavner, J. A., Karney, B. R., & Bradbury, T. N. (2016). Does couples’ communication predict marital satisfaction, or does marital satisfaction predict communication?. Journal of Marriage and Family, 78(3), 680–694.

    Open access · PMC4852543Backs: What the evidence doesn’t say The paper that argues against our premise.

  17. Karney, B. R., & Bradbury, T. N. (1995). The longitudinal course of marital quality and stability: A review of theory, method, and research. Psychological Bulletin, 118(1), 3–34.

    doi.org/10.1037/0033-2909.118.1.3Backs: What the evidence doesn’t say The three-legged model, and which leg we work on.

  18. Girme, Y. U., Overall, N. C., & Faingataa, S. (2014). “Date nights” take two: The maintenance function of shared relationship activities. Personal Relationships, 21(1), 125–149.

    doi.org/10.1111/pere.12020Backs: What the evidence doesn’t say What the weekly slot is built from, since no trial of one exists.

Reading about a pattern is a different thing from having it named while it’s happening to you. Here is how a session actually goes, and the plain-English pieces apply this to specific fights.